EDBT 2026 Demo / reviewers in the wild / expert
Chia-Hui Chang
dblp:84/3307
· DBLP profile ↗
66ranked-venue papers
22as first author
15since 2021 · last 2026
0000-0002-1101-6337ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 34 · 15 first-author · 3 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An ORID-Structured GenAI Reading Companion Integrating Structured Book Chat and Virtual Labs for Science Reading
Hui-Chun Hung, Chih-Hao Hsu, Shu-I Fang, Chen-Chung Liu, Chia-Hui Chang, Ying-Tien Wu |
AIED (3) | 6 |
| 2024 | Prosecutorial Outcome Predication with LoRA and QLoRAabstractOur study on Legal Judgment Prediction (LJP) focuses on indictments, designing innovative tasks for prosecutors to predict reasons, imprisonment, fines, and penalty types. We investigated multi-task learning (MTL), Low-Rank Adaptation (LoRA), and its quantized version (QLoRA). Results show that the MTL framework, combined with LoRA and QLoRA, excels in prediction while optimizing resource use. Error analysis in the Reason field of TWPJD dataset was conducted to guide future predictions. Future work will focus on balancing sub-task performance and collecting more investigative data to enhance legal AI applications. Kuo-Chun Chien, Chia-Hui Chang, Huai-Hsuan Huang, Jo-Chi Kung |
JURIX | 2 |
| 2024 | A Narrative Assistant for Traffic Accidents Based on Large Language Models (LLM)abstractThis study introduces the Collision Clarification Generator (CCG), a Large Language Model-based system designed to assist in documenting traffic accidents. The CCG comprises three modules: Questioning, Information Extraction, and Accident Sequence Generation, which collectively streamline the process of gathering and structuring accident information. The system employs predefined question templates and a standardized Traffic Accident Record Format (TARF) to ensure comprehensive data collection. Evaluation of the CCG involved both human assessment and LLM-based automatic evaluation. Results showed an F1 score of 0.909 in human evaluation, and scores exceeding 7 out of 10 for accuracy and completeness in LLM-based assessment. These findings demonstrate the CCG’s effectiveness in accurately documenting accident information, potentially facilitating subsequent legal and insurance processes. Jo-Chi Kung, Chia-Hui Chang, Huai-Hsuan Huang, Kuo-Chun Chien |
JURIX | 2 |
| 2024 | Legal knowledge management for prosecutors based on judgment prediction and error analysis from indictments
Kuo-Chun Chien, Chia-Hui Chang, Ren-Der Sun |
Comput. Law Secur. Rev. | 2 |
| 2024 | IoT-Based Strawberry Disease Detection With Wall-Mounted Monitoring CamerasabstractThis article proposes StrawberryTalk, an Internet of Things (IoT) platform for image-based strawberry disease detection. StrawberryTalk reuses the wall-mounted monitoring cameras without extra hardware cost. The contributions of StrawberryTalk are the utilization of IoT for automatic photo shoot and the wind detection mechanism to eliminate the obscure photos due to the wind effects. Also, the data preprocessing and the infection detection models are manipulated as IoT devices to simplify the implementation of the multicascade artificial intelligence (AI) models. We derive the relationship between the camera zoom factor and the distance between the camera and the strawberries for optimal disease detection. Accuracy of detection may be affected by obscure photos. In terms of eliminating obscure photos due to wind effect, we analytically derive the relationship between the wind alert delay and the number of obscure photos that must be retaken. For the greenhouses in the Bao Mountain, we only need to retake one photo. Based on the experiments, the mean average precision (mAP) of StrawberryTalk (to detect exact spots in a leaf) can be up to 92.37%, which is better than the previous approaches. To detect if a pot has infected leaves, the accuracy of StrawberryTalk can be up to 97.92% at the zoom factor$30\times $. In commercial operation, it is important to detect all infected strawberry pots. StrawberryTalk is able to detect all infected pots (i.e., recall is 100%) with a camera of zoom factor of$12\times $. The accuracy is 96.88%. Yi-Bing Lin, Chun-You Liu, Chia-Hui Chang, Fung-Ling Ng, Krista Yang, Jerry Hsung |
IEEE Internet Things J. | 4 |
| 2023 | Improving Chinese Fact Checking via Prompt Based Learning and Evidence RetrievalabstractVerifying the accuracy of information is a constant task as the prevalence of misinformation on the Web. In this paper, we focus on Chinese fact-checking (CHEF dataset) [1] and improve the performance through prompt-based learning in both evidence retrieval and claim verification. We adopted the Automated Prompt Engineering (APE) technique to generate the template and compared various prompt-based learning training strategies, such as prompt tuning and low-rank adaptation (LoRA) for claim verification. The research results show that prompt-based learning can improve the macro-F1 performance of claim verification by 2%-3% (from 77.62 to 80.29) using golden evidences and 110M BERT based model. For evidence retrieval, we use both the supervised SentenceBERT [2] and unsupervised PromptBERT [3] models to improve evidence retrieval performance. Experimental results show that the micro-F1 performance of evidence retrieval is significantly improved from 11.86% to 30.61% and 88.15% by PromptBERT and SentenceBERT, respectively. Finally, the overall fact-checking performance, i.e. the macro-F1 performance of claim verification, can be significantly improved from 61.94% to 80.16% when the semantic ranking-based evidence retrieval is replaced by SentenceBERT. Yu-Yen Ting, Chia-Hui Chang |
ASONAM | 2 |
| 2023 | The Design and Analysis of a Storytelling Chatbot with Natural Language Processing Techniques for Enhancing EFL ReadingabstractThis study explored the applications of natural language processing techniques on the analysis of storybook to design the storytelling chatbots. Two natural language processing techniques, co-reference and dependency parsing techniques, were applied to process the text of 157 stories books as a means to implement in the invite-narration, positive recognition, follow-up strategies when the chatbot retells the story with the students. The results of the study found that 19 (42%) students demonstrated full narratives which can re-tell the overall story in the storybook, while 22 (49%) students could only retell a simplified story. Affordances and limitations of the use of the two techniques in developing chatbots for reading are discussed. Yong Ting Feng, Chen-Chung Liu, Chia-Hui Chang |
ICALT | 3 |
| 2023 | ICCE 2023 Learning Outcomes of Computer Programming and Information Technology - Integrated Courses for Non-Computer Science Majors: Case Study of a Public Research University in TaiwanabstractThis study investigates non-CS major students' performance in computer programming and Information Technology-integrated (IT I) courses at a Taiwanese research university. Non-CS students struggle in introductory programming due to syntax and logical thinking limitations, resulting in lower grades compared to advanced programming. Similarly, ITI course grades are lower due to subject-specific demands. Gender and entry channels impact outcomes, with females excelling in informal learning. Favorable results are seen in individual applications and Multi-star Projects. Challenges include different learning paces and cultural adjustments. Regression analysis shows Introduction to Computer Programming (ICP), Advanced Computer Programming (ACP) and gender significantly affect ITI performance, explaining 24% of its variance. Recommendations include diverse teaching methods, problem-solving guidance, practical programming, collaboration, and project participation to enhance skills. Che-Yu Hsu, Feng-Nan Hwang, Tseng-Yi Chen, Chia-Hui Chang |
ICCE | 4 |
| 2023 | New Horizons of Legal Judgement Predication via Multi-Task Learning and LoRAabstractLegal Judgment Prediction (LJP) aims to predict the judgement results (such as legal article, charge and penalty) based on the criminal facts of the case. Most previous research in this field was based on criminal statements from court verdicts. However, each verdict actually is based on the content from indictments. For prosecutors, will the case be dismissed or processed? If the case is accepted, is the penalty a jail sentence or a fine? What is the charge and article violated? In this study, we therefore define three novel LJP tasks for prosecutors, including prosecution outcome prediction (LJP#1), imprison prediction (LJP#2) and fine prediction (LJP#3). We explore various multi-task learning (MTL) framework based on Word2Vec and BERT language model (LM) with either topology-based or message-passing mechanism. Moreover, we employed the LoRA (Low-Rank Adaptation) technique to save both computation time and resources during fine-tuning. Experimental results demonstrated that Word2Vec-based model combined with message passing architecture still has the potential to outperform large LM like BERT, while BERT-based models with a simple parallel architecture generally performed well. Finally, using LoRA for fine-tuning not only reduced training time (by 45%) but also improved performance (2.5% F1) in some LJP tasks. Chia-Hui Chang, Kuo-Chun Chien |
JURIX | 1 |
| 2023 | Improving colloquial case legal judgment prediction via abstractive text summarization
Yu-Xiang Hong, Chia-Hui Chang |
Comput. Law Secur. Rev. | 2 |
| 2022 | Automatic Web Data API Creation via Cross-Lingual Neural Pagination Recognition
Chia-Hui Chang, Cheng-Ju Wu, Tzu-Ping Lin |
ICWE | 1 |
| 2022 | Chat-log Disentanglement via Same-Thread Classification and Direct-Reply Prediction
Chia-Hui Chang, ZhiXian Liu, Thamolwan Poopradubsil, Yu-Ching Liao, Yu-Hao Wu |
PACLIC | 1 |
| 2022 | Question-Answer Pairing from IM Conversations via Message Merging and Reply-to Prediction
Thamolwan Poopradubsil, Chia-Hui Chang |
PACLIC | 2 |
| 2022 | Event Source Page Discovery via Policy-Based RL with Multi-task Neural Sequence Model
Chia-Hui Chang, Yu-Ching Liao, Ting Yeh |
WISE | 1 |
| 2021 | Antibody Watch: Text mining antibody specificity from the literatureabstractAntibodies are widely used reagents to test for expression of proteins and other antigens. However, they might not always reliably produce results when they do not specifically bind to the target proteins that their providers designed them for, leading to unreliable research results. While many proposals have been developed to deal with the problem of antibody specificity, it is still challenging to cover the millions of antibodies that are available to researchers. In this study, we investigate the feasibility of automatically generating alerts to users of problematic antibodies by extracting statements about antibody specificity reported in the literature. The extracted alerts can be used to construct an "Antibody Watch" knowledge base containing supporting statements of problematic antibodies. We developed a deep neural network system and tested its performance with a corpus of more than two thousand articles that reported uses of antibodies. We divided the problem into two tasks. Given an input article, the first task is to identify snippets about antibody specificity and classify if the snippets report that any antibody exhibits non-specificity, and thus is problematic. The second task is to link each of these snippets to one or more antibodies mentioned in the snippet. The experimental evaluation shows that our system can accurately perform the classification task with 0.925 weighted F1-score, linking with 0.962 accuracy, and 0.914 weighted F1 when combined to complete the joint task. We leveraged Research Resource Identifiers (RRID) to precisely identify antibodies linked to the extracted specificity snippets. The result shows that it is feasible to construct a reliable knowledge base about problematic antibodies by text mining. Chun-Nan Hsu, Chia-Hui Chang, Thamolwan Poopradubsil, Amanda Lo, Karen A. William, Ko-Wei Lin, Anita E. Bandrowski, Ibrahim Burak Özyurt, Jeffrey S. Grethe, Maryann E. Martone |
PLoS Comput. Biol. | 2 |
| 2020 | DCADE: divide and conquer alignment with dynamic encoding for full page data extraction
Oviliani Yenty Yuliana, Chia-Hui Chang |
Appl. Intell. | 2 |
| 2020 | On the Construction of Web NER Model Training Tool based on Distant SupervisionabstractNamed entity recognition (NER) is an important task in natural language understanding, as it extracts the key entities (person, organization, location, date, number, etc.) and objects (product, song, movie, activity name, etc.) mentioned in texts. However, existing natural language processing (NLP) tools (such as Stanford NER) recognize only general named entities or require annotated training examples and feature engineering for supervised model construction. Since not all languages or entities have public NER support, constructing a tool for NER model training is essential for low-resource language or entity information extraction. In this article, we study the problem of developing a tool to prepare training corpus from the Web with known seed entities for custom NER model training via distant supervision. The major challenge of automatic labeling lies in the long labeling time due to large corpus and seed entities as well as the concern to avoid false positive and false negative examples due to short and long seeds. To solve this problem, we adopt locality-sensitive hashing (LSH) for various length of seed entities. We conduct experiments on five types of entity recognition tasks, including Chinese person names, food names, locations, points of interest (POIs), and activity names to demonstrate the improvements with the proposed Web NER model construction tool. Because the training corpus is obtained by automatic labeling of the seed entity–related sentences, one could use either the entire corpus or the positive only sentences for model training. Based on the experimental results, we found the decision should depend on whether traditional linear chained conditional random fields (CRF) or deep neural network–based CRF is used for model training as well as the completeness of the provided seed list. Chien-Lung Chou, Chia-Hui Chang, Yuan-Hao Lin, Kuo-Chun Chien |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Semantic role labeling for opinion target extraction from chinese social networkabstractSocial network is a good resource to collect public opinions considering the diversity and variety in fashion, especially user generated content (UGC). Extracting the opinions from UGC can be the base of commercial policy, so how to extract the opinions correctly is an important problem. However, there are two problems: Are the opinions really talking about the target entities? Or the amount of opinions is enough for network volume analysis? In this study, we combine rule-based method and Semantic Role Labeling (SRL) to detect the opinion target (OTD) from UGC. In SRL task, we design a neural network model to extract the opinion words. Considering gradient vanish problem, we use highway connection to control the percentage of data. To improve the performance of SRL, we use additional features to help model learn the features between verb and other phrases. Finally, we design a rule to filter the result of SRL for OTD task. The experimental results show that our SRL model obtains 71% F1, and our method on OTD task obtains 73% precision which outperforms LTP. This research can be the basis of hot topic prediction for sponsors and helps them to decide marketing strategies Gui-Ru Li, Chia-Hui Chang |
ASONAM | 2 |
| 2018 | A novel alignment algorithm for effective web data extraction from singleton-item pages
Oviliani Yenty Yuliana, Chia-Hui Chang |
Appl. Intell. | 2 |
| 2017 | Implement of agent with role-based hierarchy access control for secure grouping IoTsabstractInternet of Things (IoTs) is an enabler for the intelligence appended to many central features of the modern world, such as hospitals, cities, grids, organizations, and buildings. The security and privacy are some of the major issues that prevent the wide adoption of IoTs. Its storage and computing power of these devices is often limited and their designs currently focus on ensuring functionality and largely ignore other requirements, including security and privacy concerns. Role-based Access Control in a hierarchy is an interesting research topic in computer security. Many investigations have been published in the literature with solutions involving assigning cryptographic keys to users for secure scheme in IoTs environments. The new security requirements are required by secure grouping IoTs, which are required for new applications of grouping IoTs. However, the existing secure schemes require a large number of costly arithmetic operations with large integers. The traditional scheme is difficult to implement among the IoTs devices groups with lower computation ability. In this paper, we present a solution, an implement of agent with Role-based Hierarchy Access Control for secure grouping IoTs, suitable for role-based access control in IoTs groups. The proposed scheme has promising characteristics such as high computational efficiency, little required memory and low cost implementation in each IoT device organized to be an grouping IoTs with a secure hierarchy key management. Hsing-Chung Chen, Chia-Hui Chang, Fang-Yie Leu |
CCNC | 2 |
| 2017 | Implement of agent with role-based hierarchy access control for secure grouping IoTsabstractInternet of Things (IoTs) is an enabler for the intelligence appended to many central features of the modern world, such as hospitals, cities, grids, organizations, and buildings. The security and privacy are some of the major issues that prevent the wide adoption of IoTs. Its storage and computing power of these devices is often limited and their designs currently focus on ensuring functionality and largely ignore other requirements, including security and privacy concerns. Role-based Access Control in a hierarchy is an interesting research topic in computer security. Many investigations have been published in the literature with solutions involving assigning cryptographic keys to users for secure scheme in IoTs environments. The new security requirements are required by secure grouping IoTs, which are required for new applications of grouping IoTs. However, the existing secure schemes require a large number of costly arithmetic operations with large integers. The traditional scheme is difficult to implement among the IoTs devices groups with lower computation ability. In this paper, we present a solution, an implement of agent with Role-based Hierarchy Access Control for secure grouping IoTs, suitable for role-based access control in IoTs groups. The proposed scheme has promising characteristics such as high computational efficiency, little required memory and low cost implementation in each IoT device organized to be an grouping IoTs with a secure hierarchy key management. Hsing-Chung Chen, Chia-Hui Chang, Fang-Yie Leu |
CCNC | 2 |
| 2016 | Efficient Page-Level Data Extraction via Schema Induction and Verification
Chia-Hui Chang, Tian-Sheng Chen, Ming-Chuan Chen, Jhung-Li Ding |
PAKDD (2) | 1 |
| 2016 | Enhancing POI search on maps via online address extraction and associated information segmentation
Chia-Hui Chang, Hsiu-Min Chuang, Chia-Yi Huang, Yueng-Sheng Su, Shu-Ying Li |
Appl. Intell. | 1 |
| 2016 | Enabling maps/location searches on mobile devices: constructing a POI database via focused crawling and information extractionabstractWith the popularity of mobile devices and smartphones, we have witnessed rapid growth in mobile applications and services, especially in location-based services (LBS). According to a mobile marketing survey, maps/location searches are among the most utilized services on smartphones. Points of interest (POIs), such as stores, shops, gas stations, parking lots, and bus stops, are particularly important for maps/location searches. Existing map services such as Google Maps and Wikimapia are constructed manually either professionally or with crowd sourcing. However, manual annotation is costly and limited in current POI search services. With the abundance of information on the Web, many store POIs can be extracted from the Web. In this paper, we focus on automatically constructing a POI database to enable store POI map searches. We propose techniques that are required to construct a POI database, including focused crawling, information extraction, and information retrieval techniques. We first crawl Yellow Page web sites to obtain vocabularies of store names. These vocabularies are then investigated with search engines to obtain sentences containing these store names from search snippets in order to train a store name recognition model. To extract POIs scattered across the Web, we propose a query-based crawler to find address-bearing pages that might be used to extract addresses and store names. We crawled 1.25 million distinct POI pairs scattered across the Web and implemented a POI search service via Apache Lucent’s search platform, called Solr. The experimental results demonstrate that the proposed geographical information retrieval model outperforms Wikimapia and a commercial app called ‘What’s the Number?’ Hsiu-Min Chuang, Chia-Hui Chang, Ting-Yao Kao, Chung-Ting Cheng, Ya-Yun Huang, Kuo-Pin Cheong |
Int. J. Geogr. Inf. Sci. | 2 |
| 2016 | Boosted Web Named Entity Recognition via Tri-TrainingabstractNamed entity extraction is a fundamental task for many natural language processing applications on the web. Existing studies rely on annotated training data, which is quite expensive to obtain large datasets, limiting the effectiveness of recognition. In this research, we propose a semisupervised learning approach for web named entity recognition (NER) model construction via automatic labeling and tri-training. The former utilizes structured resources containing known named entities for automatic labeling, while the latter makes use of unlabeled examples to improve the extraction performance. Since this automatically labeled training data may contain noise, a self-testing procedure is used as a follow-up to remove low-confidence annotation and prepare higher-quality training data. Furthermore, we modify tri-training for sequence labeling and derive a proper initialization for large dataset training to improve entity recognition. Finally, we apply this semisupervised learning framework for person name recognition, business organization name recognition, and location name extraction. In the task of Chinese NER, an F-measure of 0.911, 0.849, and 0.845 can be achieved, for person, business organization, and location NER, respectively. The same framework is also applied for English and Japanese business organization name recognition and obtains models with performance of a 0.832 and 0.803 F-measure. Chien-Lung Chou, Chia-Hui Chang, Ya-Yun Huang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Job-Level Alpha-Beta SearchabstractAn approach called generic job-level (JL) search was proposed to solve computer game applications by dispatching jobs to remote workers for parallel processing. This paper applies JL search to alpha-beta search, and proposes a JL alpha-beta search (JL-ABS) algorithm based on a best-first search version of MTD(f). The JL-ABS algorithm is demonstrated by using it in an opening book analysis for Chinese chess. The experimental results demonstrated that JL-ABS reached a speed-up of 10.69 when using 16 workers in the JL system. Jr-Chang Chen, I-Chen Wu, Wen-Jie Tseng 0001, Bo-Han Lin, Chia-Hui Chang |
IEEE Trans. Comput. Intell. AI Games | 5 |
| 2014 | Parallel Co-clustering with Augmented Matrices Algorithm with Map-Reduce
Meng-Lun Wu, Chia-Hui Chang |
DaWaK | 2 |
| 2014 | Enhancement of kernel dependency estimation with information generalization and a case study on skewed data
Qingzhi Chen, Chia-Hui Chang |
Appl. Intell. | 2 |
| 2014 | Integrating content-based filtering with collaborative filtering using co-clustering with augmented matrices
Meng-Lun Wu, Chia-Hui Chang, Rui-Zhe Liu |
Expert Syst. Appl. | 2 |
| 2013 | Page-Level Wrapper Verification for Unsupervised Web Data Extraction
Chia-Hui Chang, Yen-Ling Lin, Kuan-Chen Lin, Mohammed Kayed |
WISE (1) | 1 |
| 2013 | Co-clustering with augmented matrix
Meng-Lun Wu, Chia-Hui Chang, Rui-Zhe Liu |
Appl. Intell. | 2 |
| 2012 | Automatic Extraction of Blog Post from Diverse Blog PagesabstractBlog post extraction is essential for researches on blogosphere. In this paper, we address the issue of extracting blog posts from diverse blog pages, which aims at automatically and precisely finding the location of each blog post. Most of the previous researches focused on extracting main content from news pages, but the problem becomes more complex when one turns to blog pages. Our research is based on the combination of maximum scoring subsequence [11] and text-to-tag ratio [18] to develop algorithms that are suitable for blog pages. The first method that we propose is PTR Scoring, which combines postto-tag ratio with maximum scoring subsequence. The second method is CRF Scoring, which applies Conditional Random Field to train a sequence labeling model and use maximum scoring subsequence to improve the accuracy of extraction. The experimental results show that CRF Scoring achieves the best F-Measure at 91.9% compared with other methods. Chia-Hui Chang, Jhih-Ming Chen |
Web Intelligence | 1 |
| 2011 | Co-clustering with Augmented Data Matrix
Meng-Lun Wu, Chia-Hui Chang, Rui-Zhe Liu |
DaWaK | 2 |
| 2011 | Blogger-Centric Contextual Advertising
Teng-Kai Fan, Chia-Hui Chang |
Expert Syst. Appl. | 2 |
| 2010 | MapMarker: Extraction of Postal Addresses and Associated Information for General Web PagesabstractAddress information is essential for people's daily life. People often need to query addresses of unfamiliar location through Web and then use map services to mark down the location for direction purpose. Although both address information and map services are available online, they are not well combined. Users usually need to copy individual address from a Web site and paste it to another Web site with map services to locate its direction. Such copy and paste operations have to be repeated if multiple addresses are listed on a single page such as public school list or apartment list. Furthermore, associated information with individual address has to be copied and included on each marker for better comprehension. Our research is devoted to automate the above process and make the combination an easier task for users. The main techniques applied here include postal address extraction and associated information extraction. We apply sequence labeling algorithm based on Conditional Random Fields (CRFs) to train models for address extraction. Meanwhile, using the extracted addresses as landmarks, we apply pattern mining to identify the boundaries of address blocks and extract associated information with each individual address. The experimental result shows high F-score at 91% for postal address extraction and 87% accuracy for associated information extraction. Chia-Hui Chang, Shu-Ying Li |
Web Intelligence | 1 |
| 2010 | Sentiment-oriented contextual advertising
Teng-Kai Fan, Chia-Hui Chang |
Knowl. Inf. Syst. | 2 |
| 2010 | FiVaTech: Page-Level Web Data Extraction from Template PagesabstractWeb data extraction has been an important part for many Web data analysis applications. In this paper, we formulate the data extraction problem as the decoding process of page generation based on structured data and tree templates. We propose an unsupervised, page-level data extraction approach to deduce the schema and templates for each individual deep Website, which contains either singleton or multiple data records in one Webpage. FiVaTech applies tree matching, tree alignment, and mining techniques to achieve the challenging task. In experiments, FiVaTech has much higher precision than EXALG and is comparable with other record-level extraction systems like ViPER and MSE. The experiments show an encouraging result for the test pages used in many state-of-the-art Web data extraction works. Mohammed Kayed, Chia-Hui Chang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Blogger-centric contextual advertisingabstractThis paper addresses the concept of Blogger-Centric Contextual Advertising, which refers to the assignment of personal ads to any blog page, chosen in according to bloggers' interests. As blogs become a platform for expressing personal opinions, they naturally contain various kinds of statements, including facts, comments and statements about personal interests, of both a positive and negative nature. To extend the concept behind the Long Tail theory in contextual advertising, we argue that web bloggers, as the constant visitors of their own blog sites, could be potential consumers who will respond to ads on their own blogs. Hence, in this paper, we propose using text mining techniques to discover bloggers' immediate personal interests in order to improve online contextual advertising. The proposed BCCA (Blogger-Centric Contextual Advertising) framework aims to combine contextual advertising matching with text mining in order to select ads that are related to personal interests as revealed in a blog and rank them according to their relevance. We validate our approach experimentally using a set of data that includes both real ads and actual blog pages. The results indicate that our proposed method could effectively identify those ads that are positively-correlated with a blogger's personal interests. Teng-Kai Fan, Chia-Hui Chang |
CIKM | 2 |
| 2009 | Sentiment-Oriented Contextual Advertising
Teng-Kai Fan, Chia-Hui Chang |
ECIR | 2 |
| 2008 | Exploring Evolutionary Technical Trends from Academic Research PapersabstractAutomatic Term Recognition (ATR) is concerned with discovering terminology in large volumes of text corpora. Technical terms are vital elements for understanding the techniques used in academic research papers, and in this paper, we use focused technical terms to explore technical trends in the research literature. The major purpose of this work is to understand the relationship between techniques and research topics to better explore technical trends. We define this new text mining issue and apply machine learning algorithms for solving this problem by (1) recognizing focused technical terms from research papers; (2) classifying these terms into predefined technology categories; (3) analyzing the evolution of technical trends. The dataset consists of 656 papers collected from well-known conferences on ACM. The experimental results indicate that our proposed methods can effectively explore interesting evolutionary technical trends in various research topics. Teng-Kai Fan, Chia-Hui Chang |
Document Analysis Systems | 2 |
| 2008 | Efficient mining of frequent episodes from complex sequences
Kuo-Yu Huang, Chia-Hui Chang |
Inf. Syst. | 2 |
| 2007 | Efficient text chunking using linear kernel with masked method
Yu-Chieh Wu, Chia-Hui Chang |
Knowl. Based Syst. | 2 |
| 2006 | Efficient Mining Strategy for Frequent Serial Episodes in Temporal Database
Kuo-Yu Huang, Chia-Hui Chang |
APWeb | 2 |
| 2006 | A General and Multi-lingual Phrase Chunking Model Based on Masking Method
Yu-Chieh Wu, Chia-Hui Chang, Yue-Shi Lee |
CICLing | 2 |
| 2006 | COBRA: Closed Sequential Pattern Mining Using Bi-phase Reduction Approach
Kuo-Yu Huang, Chia-Hui Chang, Jiun-Hung Tung, Cheng-Tao Ho |
DaWaK | 2 |
| 2006 | A Survey of Web Information Extraction SystemsabstractThe Internet presents a huge amount of useful information which is usually formatted for its users, which makes it difficult to extract relevant data from various sources. Therefore, the availability of robust, flexible Information Extraction (IE) systems that transform the Web pages into program-friendly structures such as a relational database will become a great necessity. Although many approaches for data extraction from Web pages have been developed, there has been limited effort to compare such tools. Unfortunately, in only a few cases can the results generated by distinct tools be directly compared since the addressed extraction tasks are different. This paper surveys the major Web data extraction approaches and compares them in three dimensions: the task domain, the automation degree, and the techniques used. The criteria of the first dimension explain why an IE system fails to handle some Web sites of particular structures. The criteria of the second dimension classify IE systems based on the techniques used. The criteria of the third dimension measure the degree of automation for IE systems. We believe these criteria provide qualitatively measures to evaluate various IE approaches. Chia-Hui Chang, Mohammed Kayed, Moheb R. Girgis, Khaled Shaalan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | ClosedPROWL: Efficient Mining of Closed Frequent Continuities by Projected Window List TechnologyabstractMining frequent patterns in databases is a fundamental and essential problem in data mining research. A continuity is a kind of causal relationship which describes a definite temporal factor with exact position between the records. Since continuities break the boundaries of records, the number of potential patterns will increase drastically. An alternative approach is to mine closed frequent continuities. Mining closed frequent patterns has the same power as mining the complete set of frequent patterns, while substantially reducing redundant rules to be generated and increasing the effectiveness of mining. In this paper, we propose a method called projected window list technology for the mining of frequent continuities. We present a closed frequent continuity mining algorithm, ClosedPROWL. Experimental result shows that our algorithm is more efficient than previously proposed algorithms. Kuo-Yu Huang, Chia-Hui Chang, Kuo-Zui Lin |
SDM | 2 |
| 2005 | Integrating Web Information to Generate Chinese Video Summaries
Yue-Shi Lee, Yu-Chieh Wu, Chia-Hui Chang |
SEKE | 3 |
| 2005 | Categorical data visualization and clustering using subjective factors
Chia-Hui Chang, Zhi-Kai Ding |
Data Knowl. Eng. | 1 |
| 2005 | Reconfigurable Web wrapper agents for biological information integrationabstractAbstract A variety of biological data is transferred and exchanged in overwhelming volumes on the World Wide Web. How to rapidly capture, utilize, and integrate the information on the Internet to discover valuable biological knowledge is one of the most critical issues in bioinformatics. Many information integration systems have been proposed for integrating biological data. These systems usually rely on an intermediate software layer called wrappers to access connected information sources. Wrapper construction for Web data sources is often specially hand coded to accommodate the differences between each Web site. However, programming a Web wrapper requires substantial programming skill, and is time‐consuming and hard to maintain. In this article we provide a solution for rapidly building software agents that can serve as Web wrappers for biological information integration. We define an XML‐based language called Web Navigation Description Language (WNDL), to model a Web‐browsing session. A WNDL script describes how to locate the data, extract the data, and combine the data. By executing different WNDL scripts, we can automate virtually all types of Web‐browsing sessions. We also describe IEPAD (Information Extraction Based on Pattern Discovery), a data extractor based on pattern discovery techniques. IEPAD allows our software agents to automatically discover the extraction rules to extract the contents of a structurally formatted Web page. With a programming‐by‐example authoring tool, a user can generate a complete Web wrapper agent by browsing the target Web sites. We built a variety of biological applications to demonstrate the feasibility of our approach. Chun-Nan Hsu, Chia-Hui Chang, Chang-Huain Hsieh, Jiann-Jyh Lu, Chien-Chi Chang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2005 | SMCA: A General Model for Mining Asynchronous Periodic Patterns in Temporal DatabasesabstractMining periodic patterns in time series databases is an important data mining problem with many applications. Previous studies have considered synchronous periodic patterns where misaligned occurrences are not allowed. However, asynchronous periodic pattern mining has received less attention and only been discussed for a sequence of symbols where each time point contains one event. In this paper, we propose a more general model of asynchronous periodic patterns from a sequence of symbol sets where a time slot can contain multiple events. Three parameters min/spl I.bar/rep, max/spl I.bar/dis, and global/spl I.bar/rep are employed to specify the minimum number of repetitions required for a valid segment of nondisrupted pattern occurrences, the maximum allowed disturbance between two successive valid segments, and the total repetitions required for a valid sequence. A 4-phase algorithm is devised to discover periodic patterns from a time series database presented in vertical format. The experiments demonstrate good performance and scalability with large frequent patterns. Kuo-Yu Huang, Chia-Hui Chang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2004 | Categorical Data Visualization and Clustering Using Subjective Factors
Chia-Hui Chang, Zhi-Kai Ding |
DaWaK | 1 |
| 2004 | Mining Periodic Patterns in Sequence Data
Kuo-Yu Huang, Chia-Hui Chang |
DaWaK | 2 |
| 2004 | PROWL: An Efficient Frequent continuity Mining Algorithm on Event Sequences
Kuo-Yu Huang, Chia-Hui Chang, Kuo-Zui Lin |
DaWaK | 2 |
| 2004 | COCOA: Compressed Continuity Analysis for Temporal Databases
Kuo-Yu Huang, Chia-Hui Chang, Kuo-Zui Lin |
PKDD | 2 |
| 2004 | Performance evaluation of a database of repetitive elements in complete genomes
Jorng-Tzong Horng, Feng-Mao Lin, Li-Cheng Wu, Chia-Hui Chang |
J. Syst. Softw. | 4 |
| 2003 | Enhancing SWF for Incremental Association Mining by Itemset Maintenance
Chia-Hui Chang, Shi-Hsan Yang |
PAKDD | 1 |
| 2003 | Automatic information extraction from semi-structured Web pages by pattern discovery
Chia-Hui Chang, Chun-Nan Hsu, Shao-Chen Lui |
Decis. Support Syst. | 1 |
| 2002 | Automatic Information Extraction for Multiple Singular Web Pages
Chia-Hui Chang, Shih-Chien Kuo, Kuo-Yu Huang, Tsung-Hsin Ho, Chih-Lung Lin |
PAKDD | 1 |
| 2001 | Applying Pattern Mining to Web Information Extraction
Chia-Hui Chang, Shao-Chen Lui, Yen-Chin Wu |
PAKDD | 1 |
| 2001 | IEPAD: information extraction based on pattern discoveryabstractArticle IEPAD: information extraction based on pattern discovery Authors: Chia-Hui Chang Dept. of Computer Science and Information Engineering, National Central University, Chung-Li, Taiwan 320 Dept. of Computer Science and Information Engineering, National Central University, Chung-Li, Taiwan 320View Profile , Shao-Chen Lui Dept. of Computer Science and Information Engineering, National Central University, Chung-Li, Taiwan 320 Dept. of Computer Science and Information Engineering, National Central University, Chung-Li, Taiwan 320View Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001 Pages 681–688https://doi.org/10.1145/371920.372182Published:01 April 2001Publication History 281citation2,397DownloadsMetricsTotal Citations281Total Downloads2,397Last 12 Months52Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Chia-Hui Chang, Shao-Chen Lui |
WWW | 1 |
| 1999 | Enabling Concept-Based Relevance Feedback for Information Retrieval on the WWWabstractThe World Wide Web is a world of great richness, but finding information on the Web is also a great challenge. Keyword-based querying has been an immediate and efficient way to specify and retrieve related information that the user inquires. However, conventional document ranking based on an automatic assessment of document relevance to the query may not be the best approach when little information is given, as in most cases. In order to clarify the ambiguity of the short queries given by users, we propose the idea of concept-based relevance feedback for Web information retrieval. The idea is to have users give two to three times more feedback in the same amount of time that would be required to give feedback for conventional feedback mechanisms. Under this design principle, we apply clustering techniques to the initial search results to provide concept-based browsing. We show the performance of various feedback interface designs and compare their pros and cons. We measure precision and relative recall to show how clustering improves performance over conventional similarity ranking and, most importantly, we show how the assistance of concept-based presentation reduces browsing labor. Chia-Hui Chang, Ching-Chi Hsu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1998 | Exploiting hyperlinks for automatic information discovery on the WWWabstractThe explosion of the World Wide Web as a global information network brings with it a number of related challenges for information retrieval and automation. The link structure, which is the main feature of the hypermedia environment, can be a rich source of information for exploration. This paper is centered around the exploitation of hyperlinks in the subject of automatic discovery. The paper discusses the design of site traversal algorithms and identifies some important parameters for hyperlink-based information discovery on the Web. Meanwhile, it proposes a new evaluation method, the search depth, for comparing various hyperlink traversal algorithms. Chia-Hui Chang, Ching-Chi Hsu, Cheng-Lin Hou |
ICTAI | 1 |
| 1998 | Constructing Personalized Information Agents
Chia-Hui Chang, Ching-Chi Hsu |
PAKDD | 1 |
| 1998 | Integrating Query Expansion and Conceptual Relevance Feedback for Personalized Web Information Retrieval
Chia-Hui Chang, Ching-Chi Hsu |
Comput. Networks | 1 |
| 1997 | Customizable Multi-Engine Search Tool with Clustering
Chia-Hui Chang, Ching-Chi Hsu |
Comput. Networks | 1 |