Pu-Jen Cheng

dblp:45/160 · DBLP profile ↗
← Back
34ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-5892-0385ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
abstract
Synthetic data augmentation via Large Language Models (LLMs) allows researchers to leverage additional training data, thus enhancing the performance of downstream tasks, especially when real-world data is scarce. However, the generated data can deviate from the real-world data, and this misalignment can bring about deficient results while applying the trained model to applications. Therefore, we proposed efficient weighted-loss approaches to align synthetic data with real-world distribution by emphasizing high-quality and diversified data generated by LLMs using merely a tiny amount of real-world data. We empirically assessed the effectiveness of our methods on multiple text classification tasks, and the results showed that leveraging our approaches on a BERT-level model robustly outperformed standard cross-entropy and other data weighting approaches, providing potential solutions to effectively leveraging synthetic data from any suitable data generator.
Hsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma, Pu-Jen Cheng
ICLR5
2024 Multi-DSI: Non-deterministic Identifier and Concept Alignment for Differentiable Search Index
abstract
With the advent of generative deep learning models, generative IR has gained increasing attention. However, existing methods face two issues: (1) when a document is represented by a single semantic ID, the retrieval model may fail to capture the multifaceted and complex content of the document; and (2) when the generated training data exhibits semantic ambiguity, the retrieval model may struggle to distinguish the differences in the content of similar documents. To address these issues, we propose Multi-DSI to (1) offer multiple non-deterministic semantic identifiers and (2) align the concepts of queries and documents to avoid ambiguity. Extensive experiments on two benchmark datasets demonstrate that Multi-DSI significantly outperforms baseline methods by 7.4%.
Yu-Ze Liu, Jyun-Yu Jiang, Pu-Jen Cheng
CIKM3
2023 Incorporating Co-purchase Correlation for Next-basket Recommendation
abstract
Next-basket recommendation (NBR) aims to recommend a set of items that users would most likely purchase together. Existing approaches use deep learning to capture basket-level preference and traditional statistical methods to model user behavior sequences. However, these methods neglect the correlation of co-purchase items among users. We, therefore, propose a novel model that incorporates Co-purchase Correlation with Bidirectional Transformer (CCBT) to enhance item representation by exploiting the correlation among users' baskets. The results of experiments conducted on four real-world datasets demonstrate the proposed model outperforms state-of-the-art NBR methods. The relative improvement for Recall@20 ranges from 11% to 27%.
Yu Hao Chou, Pu-Jen Cheng
CIKM2
2023 printf: Preference Modeling Based on User Reviews with Item Images and Textual Information via Graph Learning
abstract
Nowadays, modern recommender systems usually leverage textual and visual contents as auxiliary information to predict user preference. For textual information, review texts are one of the most popular contents to model user behaviors. Nevertheless, reviews usually lose their shine when it comes to top-N recommender systems because those that solely utilize textual reviews as features struggle to adequately capture the interaction relationships between users and items. For visual one, it is usually modeled with naive convolutional networks and also hard to capture high-order relationships between users and items. Moreover, previous works did not collaboratively use both texts and images in a proper way. In this paper, we propose printf, preference modeling based on user reviews with item images and textual information via graph learning, to address the above challenges. Specifically, the dimension-based attention mechanism directs relations between user reviews and interacted items, allowing each dimension to contribute different importance weights to derive user representations. Extensive experiments are conducted on three publicly available datasets. The experimental results demonstrate that our proposed printf consistently outperforms baseline methods with the relative improvements for NDCG@5 of 26.80%, 48.65%, and 25.74% on Amazon-Grocery, Amazon-Tools, and Amazon-Electronics datasets, respectively. The in-depth analysis also indicates the dimensions of review representations definitely have different topics and aspects, assisting the validity of our model design.
Hao-Lun Lin, Jyun-Yu Jiang, Ming-Hao Juan, Pu-Jen Cheng
CIKM4
2023 Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs
abstract
Aiming to produce reinforcement learning (RL) policies that are human-interpretable and can generalize better to novel scenarios, Trivedi et al. (2021) present a method (LEAPS) that first learns a program embedding space to continuously parameterize diverse programs from a pre-generated program dataset, and then searches for a task-solving program in the learned program embedding space when given a task. Despite the encouraging results, the program policies that LEAPS can produce are limited by the distribution of the program dataset. Furthermore, during searching, LEAPS evaluates each candidate program solely based on its return, failing to precisely reward correct parts of programs and penalize incorrect parts. To address these issues, we propose to learn a meta-policy that composes a series of programs sampled from the learned program embedding space. By learning to compose programs, our proposed hierarchical programmatic reinforcement learning (HPRL) framework can produce program policies that describe out-of-distributionally complex behaviors and directly assign credits to programs that induce desired behaviors. The experimental results in the Karel domain show that our proposed framework outperforms baselines. The ablation studies confirm the limitations of LEAPS and justify our design choices.
Guan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-yi Lee, Shao-Hua Sun
ICML3
2022 R-TeaFor: Regularized Teacher-Forcing for Abstractive Summarization
abstract
Teacher-forcing is widely used in training sequence generation models to improve sampling efficiency and to stabilize training.However, teacher-forcing is vulnerable to the exposure bias problem.Previous works have attempted to address exposure bias by modifying the training data to simulate model-generated results.Nevertheless, they do not consider the pairwise relationship between the original training data and the modified ones, which provides more information during training.Hence, we propose Regularized Teacher-Forcing (R-TeaFor) to utilize this relationship for better regularization.Empirically, our experiments show that R-TeaFor outperforms previous summarization state-of-the-art models, and the results can be generalized to different pre-trained models.
Guan-Yu Lin, Pu-Jen Cheng
EMNLP2
2022 Sub-Resolution Assist Feature Generation with Reinforcement Learning and Transfer Learning
abstract
As modern photolithography feature sizes continue to shrink, sub-resolution assist feature (SRAF) generation has become a key resolution enhancement technique to improve the manufacturing process window. State-of-the-art works resort to machine learning to overcome the deficiencies of model-based and rule-based approaches. Nevertheless, these machine learning-based methods do not consider or implicitly consider the optical interference between SRAFs, and highly rely on post-processing to satisfy SRAF mask manufacturing rules. In this paper, we are the first to generate SRAFs using reinforcement learning to address SRAF interference and produce mask-rule-compliant results directly. In this way, our two-phase learning enables us to emulate the style of model-based SRAFs while further improving the process variation (PV) band. A state alignment and action transformation mechanism is proposed to achieve orientation equivariance while expediting the training process. We also propose a transfer learning framework, allowing SRAF generation under different light sources without retraining the model. Compared with state-of-the-art works, our method improves the solution quality in terms of PV band and edge placement error (EPE) while reducing the overall runtime.
Guan-Ting Liu, Wei-Chen Tai, Iris Hui-Ru Jiang, James P. Shiely, Pu-Jen Cheng
ICCAD6
2021 End-to-End Video Question-Answer Generation With Generator-Pretester Network
abstract
We study a novel task, Video Question-Answer Generation (VQAG), for challenging Video Question Answering (Video QA) task in multimedia. Due to expensive data annotation costs, many widely used, large-scale Video QA datasets such as Video-QA, MSVD-QA and MSRVTT-QA are automatically annotated using Caption Question Generation (CapQG) which inputs captions instead of the video itself. As captions neither fully represent a video, nor are they always practically available, it is crucial to generate question-answer pairs based on a video via Video Question-Answer Generation (VQAG). Existing video-to-text (V2T) approaches, despite taking a video as the input, only generate a question alone. In this work, we propose a novel model Generator-Pretester Network that focuses on two components: (1) The Joint Question-Answer Generator (JQAG) which generates a question with its corresponding answer to allow Video Question “Answering” training. (2) The Pretester (PT) verifies a generated question by trying to answer it and checks the pretested answer with both the model’s proposed answer and the ground truth answer. We evaluate our system with the only two available large-scale human-annotated Video QA datasets and achieves state-of-the-art question generation performances. Furthermore, using our generated QA pairs only on the Video QA task, we can surpass some supervised baselines. As a pre-training strategy, we outperform both CapQG and transfer learning approaches when employing semi-supervised (20%) or fully supervised learning with annotated data. These experimental results suggest the novel perspectives for Video QA training.
Hung-Ting Su, Chen-Hsi Chang, Po-Wei Shen, Yu-Siang Wang, Ya-Liang Chang, Pu-Jen Cheng, Winston H. Hsu
IEEE Trans. Circuits Syst. Video Technol.7
2020 Improving One-class Recommendation with Multi-tasking on Various Preference Intensities
abstract
In the one-class recommendation problem, it’s required to make recommendations basing on users’ implicit feedback, which is inferred from their action and inaction. Existing works obtain representations of users and items by encoding positive and negative interactions observed from training data. However, these efforts assume that all positive signals from implicit feedback reflect a fixed preference intensity, which is not realistic. Consequently, representations learned with these methods usually fail to capture informative entity features that reflect various preference intensities.
Chu-Jen Shao, Hao-Ming Fu, Pu-Jen Cheng
RecSys3
2019 Learning Unsupervised Semantic Document Representation for Fine-grained Aspect-based Sentiment Analysis
abstract
Document representation is the core of many NLP tasks on machine understanding. A general representation learned in an unsupervised manner reserves generality and can be used for various applications. In practice, sentiment analysis (SA) has been a challenging task that is regarded to be deeply semantic-related and is often used to assess general representations. Existing methods on unsupervised document representation learning can be separated into two families: sequential ones, which explicitly take the ordering of words into consideration, and non-sequential ones, which do not explicitly do so. However, both of them suffer from their own weaknesses. In this paper, we propose a model that overcomes difficulties encountered by both families of methods. Experiments show that our model outperforms state-of-the-art methods on popular SA datasets and a fine-grained aspect-based SA by a large margin.
Hao-Ming Fu, Pu-Jen Cheng
SIGIR2
2019 Cocluster hypothesis and ranking consistency for relevance ranking in web search
abstract
Conventional approaches to relevance ranking typically optimize ranking models by each query separately. The traditional cluster hypothesis also does not consider the dependency between related queries. The goal of this paper is to leverage similar search intents to perform ranking consistency so that the search performance can be improved accordingly. Different from the previous supervised approach, which learns relevance by click‐through data, we propose a novel cocluster hypothesis to bridge the gap between relevance ranking and ranking consistency. A nearest‐neighbors test is also designed to measure the extent to which the cocluster hypothesis holds. Based on the hypothesis, we further propose a two‐stage unsupervised approach, in which two ranking heuristics and a cost function are developed to optimize the combination of consistency and uniqueness (or inconsistency). Extensive experiments have been conducted on a real and large‐scale search engine log. The experimental results not only verify the applicability of the proposed cocluster hypothesis but also show that our approach is effective in boosting the retrieval performance of the commercial search engine and reaches a comparable performance to the supervised approach.
Jian-De Jiang, Jyun-Yu Jiang, Pu-Jen Cheng
J. Assoc. Inf. Sci. Technol.3
2018 Translating Representations of Knowledge Graphs with Neighbors
abstract
Knowledge graph completion is a critical issue because many applications benefit from their structural and rich resources. In this paper, we propose a method named TransN, which consid- ers the dependencies between triples and incorporates neighbor information dynamically. In experiments, we evaluate our model by link prediction and also conduct several qualitative analyses to prove effectiveness. Experimental results show that our model could integrate neighbor information effectively and outperform state-of-the-art models.
Chun-Chih Wang, Pu-Jen Cheng
SIGIR2
2017 Improving One-Class Collaborative Filtering with Manifold Regularization by Data-driven Feature Representation
Yen-Chieh Lien, Pu-Jen Cheng
PAKDD (2)2
2017 Open Source Repository Recommendation in Social Coding
abstract
Social coding and open source repositories have become more and more popular. Software developers have various alternatives to contribute themselves to the communities and collaborate with others. However, nowadays there is no effective recommender suggesting developers appropriate repositories in both the academia and the industry. Although existing one-class collaborative filtering (OCCF) approaches can be applied to this problem, they do not consider particular constraints of social coding such as the programming languages, which, to some extent, associate the repositories with the developers. The aim of this paper is to investigate the feasibility of leveraging user programming language preference to improve the performance of OCCF-based repository recommendation. Based on matrix factorization, we propose language-regularized matrix factorization (LRMF), which is regularized by the relationships between user programming language preferences. Extensive experiments have been conducted on the real-world dataset of GitHub. The results demonstrate that our framework significantly outperforms five competitive baselines.
Jyun-Yu Jiang, Pu-Jen Cheng, Wei Wang 0010
SIGIR2
2016 Dynamically Integrating Item Exposure with Rating Prediction in Collaborative Filtering
abstract
The paper proposes a novel approach to appropriately promote those items with few ratings in collaborative filtering. Different from previous works, we force the items with few ratings to be promoted to the users who would potentially be able to give ratings, and then leverage the gathered user preference to punish the promoted items with low quality intrinsically. By slightly sacrificing the benefit of recommending the best items in terms of user satisfaction, our approach seeks to provide all of the items with a chance to be visible equally. The results of the experiments conducted on MovieLens and Netflix data demonstrate its feasibility.
Ting-Yi Shih, Ting-Chang Hou, Jian-De Jiang, Yen-Chieh Lien, Chia-Rui Lin, Pu-Jen Cheng
SIGIR6
2015 Improving Ranking Consistency for Web Search by Leveraging a Knowledge Base and Search Logs
abstract
In this paper, we propose a new idea called ranking consistency in web search. Relevance ranking is one of the biggest problems in creating an effective web search system. Given some queries with similar search intents, conventional approaches typically only optimize ranking models by each query separately. Hence, there are inconsistent rankings in modern search engines. It is expected that the search results of different queries with similar search intents should preserve ranking consistency. The aim of this paper is to learn consistent rankings in search results for improving the relevance ranking in web search. We then propose a re-ranking model aiming to simultaneously improve relevance ranking and ranking consistency by leveraging knowledge bases and search logs. To the best of our knowledge, our work offers the first solution to improving relevance rankings with ranking consistency. Extensive experiments have been conducted using the Freebase knowledge base and the large-scale query-log of a commercial search engine. The experimental results show that our approach significantly improves relevance ranking and ranking consistency. Two user surveys on Amazon Mechanical Turk also show that users are sensitive and prefer the consistent ranking results generated by our model.
Jyun-Yu Jiang, Jing Liu 0022, Chin-Yew Lin, Pu-Jen Cheng
CIKM4
2015 An efficient multiple session key establishment scheme for VANET group integration
abstract
VANET (Vehicular Ad-hoc Network) is the one mainly utilized to create communication networks for vehicles or other roadside devices so that they can quickly share and receive messages. Nevertheless, VANET belongs to the family of wireless networks, which means VANET functions are unsafe. In order to provide safe communication channels, we introduce the key agreements technology to VANET communication. Traditional key agreement schemes, however, are inefficient and would consume too much of the resources especially when they are handling large groups of users or when groups are to be combined. To improve the efficiency, we use the Chinese remainder theorem to build a batch key agreement protocol instead. The improved key for VANET environments is a safer and quicker way to establish communication channels.
Cheng-Chi Lee, Yan-Ming Lai, Pu-Jen Cheng
Intelligent Vehicles Symposium3
2015 Semantic Tagging of Mathematical Expressions
abstract
Semantic tagging of mathematical expressions (STME) gives semantic meanings to tokens in mathematical expressions. In this work, we propose a novel STME approach that relies on neither text along with expressions, nor labelled training data. Instead, our method only requires a mathematical grammar set. We point out that, besides the grammar of mathematics, the special property of variables and user habits of writing expressions help us understand the implicit intents of the user. We build a system that considers both restrictions from the grammar and variable properties, and then apply an unsupervised method to our probabilistic model to learn the user habits. To evaluate our system, we build large-scale training and test datasets automatically from a public math forum. The results demonstrate the significant improvement of our method, compared to the maximum-frequency baseline. We also create statistics to reveal the properties of mathematics language.
Pao-Yu Chien, Pu-Jen Cheng
WWW2
2014 Learning user reformulation behavior for query auto-completion
abstract
It is crucial for query auto-completion to accurately predict what a user is typing. Given a query prefix and its context (e.g., previous queries), conventional context-aware approaches often produce relevant queries to the context. The purpose of this paper is to investigate the feasibility of exploiting the context to learn user reformulation behavior for boosting prediction performance. We first conduct an in-depth analysis of how the users reformulate their queries. Based on the analysis, we propose a supervised approach to query auto-completion, where three kinds of reformulation-related features are considered, including term-level, query-level and session-level features. These features carefully capture how the users change preceding queries along the query sessions. Extensive experiments have been conducted on the large-scale query log of a commercial search engine. The experimental results demonstrate a significant improvement over 4 competitive baselines.
Jyun-Yu Jiang, Yen-Yu Ke, Pao-Yu Chien, Pu-Jen Cheng
SIGIR4
2013 The Impact of Social Diversity and Dynamic Influence Propagation for Identifying Influencers in Social Networks
abstract
There has been significant recent interest in using the aggregate information from social media sites (e.g., Twitter) to identify influencers. To investigate this issue, one dynamic diversity-dependent algorithm is proposed for detecting the influencers by evaluating the influence of users throughout social networks. Comparative analyses with the existing methods on either synthetic social networks or real Twitter data show that our strategy performs best. It implies that the pattern of the influence propagation should be updated dynamically to reflect the flow of influence spread to better capture the rapidly changing dynamics of microblogs. Our proposed scheme is therefore practical and feasible to be deployed in the real world.
Pei-Ying Huang, Hsin-Yu Liu, Chin-Hui Chen, Pu-Jen Cheng
Web Intelligence4
2012 Visualizing timelines: evolutionary summarization via iterative reinforcement between text and image streams
abstract
We present a novel graph-based framework for timeline summarization, the task of creating different summaries for different timestamps but for the same topic. Our work extends timeline summarization to a multimodal setting and creates timelines that are both textual and visual. Our approach exploits the fact that news documents are often accompanied by pictures and the two share some common content. Our model optimizes local summary creation and global timeline generation jointly following an iterative approach based on mutual reinforcement and co-ranking. In our algorithm, individual summaries are generated by taking into account the mutual dependencies between sentences and images, and are iteratively refined by considering how they contribute to the global timeline and its coherence. Experiments on real-world datasets show that the timelines produced by our model outperform several competitive baselines both in terms of ROUGE and when assessed by human evaluators.
Rui Yan 0001, Xiaojun Wan 0001, Mirella Lapata, Wayne Xin Zhao, Pu-Jen Cheng, Xiaoming Li 0001
CIKM5
2012 Learning-based time-sensitive re-ranking for web search
abstract
To model time-dependent user intent for Web search, this paper proposes a novel method using machine learning techniques to exploit temporal features for effective time-sensitive search result re-ranking. We propose models to incorporate users' click through information for queries that are seen in the training data, and then further extend the model to deal with unseen queries considering the relationship between queries. Experiment shows significant improvement on search result ranking over original search outputs.
Po-Tzu Chang, Yen-Chieh Huang, Cheng-Lun Yang, Shou-De Lin, Pu-Jen Cheng
SIGIR5
2012 Event Duration Detection on Microblogging
abstract
Microblog users often post what they observe in the surroundings, making it possible to use such microblog data to perform event detection. In this paper, we propose an issue of event duration detection and choose rain as our target event. Our goal is to construct an online virtual weather station, which reports local weather condition such as rain with the microblog data. Different from previous work focusing on earthquake, rain is a relatively minor event and may continue for a period of time. The virtual station, therefore, needs to detect not only when it starts but also when it finishes. The system trains a classifier to extract the event of interest. A wavelet-based method and aging-based method are proposed to detect the beginning of the rain event and its duration, respectively. Our experiments are conducted on real data, collected from Twitter and online weather stations. The results of experiments show the feasibility of virtual weather system. Our user behavior analysis also explains why the system works.
Yi-Shiang Tzeng, Jyun-Yu Jiang, Pu-Jen Cheng
Web Intelligence3
2011 Query sampling for learning data fusion
abstract
Data fusion is to merge the results of multiple independent retrieval models into a single ranked list. Several earlier studies have shown that the combination of different models can improve the retrieval performance better than using any of the individual models. Although many promising results have been given by supervised fusion methods, training data sampling has attracted little attention in previous work of data fusion. By observing some evaluations on TREC and NTCIR datasets, we found that the performance of one model varied largely from one training example to another, so that not all training examples were equivalently effective. In this paper, we propose two novel approaches: greedy and boosting approaches, which select effective training data by query sampling to improve the performance of supervised data fusion algorithms such as BayesFuse, probFuse and MAPFuse. Extensive experiments were conducted on five data sets including TREC-3,4,5 and NTCIR-3,4. The results show that our sampling approaches can significantly improve the retrieval performance of those data fusion methods.
Ting-Chu Lin, Pu-Jen Cheng
CIKM2
2011 Clustering and Visualizing Geographic Data Using Geo-tree
abstract
Plotting lots of geographical data points usually clutters up a map. In this paper, we propose an approach to provide a summary view of geographical data by efficiently clustering. We present a novel data structure, called Geo-tree, which is extended from quad tree, and then develop two algorithms, which use Geo-tree to cluster geographic data and visualize the clusters with a heat map-like representation. The experimental results show that our approach is very efficient in a large scale, compared to K-means and HAC, and the clustering results are comparable to theirs.
Che-An Lu, Chin-Hui Chen, Pu-Jen Cheng
Web Intelligence3
2010 To translate or not to translate?
abstract
Query translation is an important task in cross-language information retrieval (CLIR) aiming to translate queries into languages used in documents. The purpose of this paper is to investigate the necessity of translating query terms, which might differ from one term to another. Some untranslated terms cause irreparable performance drop while others do not. We propose an approach to estimate the translation probability of a query term, which helps decide if it should be translated or not. The approach learns regression and classification models based on a rich set of linguistic and statistical properties of the term. Experiments on NTCIR-4 and NTCIR-5 English-Chinese CLIR tasks demonstrate that the proposed approach can significantly improve CLIR performance. An in-depth analysis is also provided for discussing the impact of untranslated out-of-vocabulary (OOV) query terms and translation quality of non-OOV query terms on CLIR performance.
Chin-Hui Chen, Shao-Hang Kao, Pu-Jen Cheng
SIGIR4
2009 Learning to rank from Bayesian decision inference
abstract
Ranking is a key problem in many information retrieval (IR) applications, such as document retrieval and collaborative filtering. In this paper, we address the issue of learning to rank in document retrieval. Learning-based methods, such as RankNet, RankSVM, and RankBoost, try to create ranking functions automatically by using some training data. Recently, several learning to rank methods have been proposed to directly optimize the performance of IR applications in terms of various evaluation measures. They undoubtedly provide statistically significant improvements over conventional methods; however, from the viewpoint of decision-making, most of them do not minimize the Bayes risk of the IR system. In an attempt to fill this research gap, we propose a novel framework that directly optimizes the Bayes risk related to the ranking accuracy in terms of the IR evaluation measures. The results of experiments on the LETOR collections demonstrate that the framework outperforms several existing methods in most cases.
Jen-Wei Kuo, Pu-Jen Cheng, Hsin-Min Wang
CIKM2
2009 A term dependency-based approach for query terms ranking
abstract
Formulating appropriate and effective queries has been regarded as a challenging issue, since a large number of candidate words or phrases could be chosen as query terms to convey users' information needs. In this paper, we propose an approach to rank a set of given query terms according their effectiveness, wherein top ranked terms will be selected as an effective query. Our ranking approach exploits and benefits from the underlying relationship between the query terms, and thereby the effective terms can be properly combined into the query. Two regression models which capture a rich set of linguistic and statistical properties are used in our approach. Experiments on NTCIR-4 ad-hoc retrieval tasks demonstrate that the proposed approach can significantly improve retrieval performance, and can be well applied to other problems such as query expansion and querying by text segments.
Ruey-Cheng Chen, Shao-Hang Kao, Pu-Jen Cheng
CIKM4
2007 LiveConcept: Web Search using Structured Query
abstract
In this paper, we propose an approach that structure the given query as a taxonomy, called a query taxonomy, from the user's perspective. The proposed approach is based on an unsupervised classification method, which uses the dynamic Web as the training corpus. With query taxonomy, users can browse relevant Web documents more conveniently and comprehensibly. Based on the proposed approach, we empirically build a search engine system, LiveConcept, and verify the feasibility and the effectiveness of the system after a series of user study.
Pu-Jen Cheng, Ching-Hsiang Tsai, Chen-Ming Hung
ICDE1
2006 Query taxonomy generation for web search
abstract
We propose an approach that organizes the search-result clusters into a hierarchical structure, called a query taxonomy, from the user's perspective. The proposed approach is based on an unsupervised classification method, which uses the dynamic Web as the training corpus. With query taxonomy, users can browse relevant Web documents more conveniently and comprehensibly. Our experimental results verify the feasibility and the effectiveness of the proposed approach to query taxonomy generation in Web search.
Pu-Jen Cheng, Ching-Hsiang Tsai, Chen-Ming Hung, Lee-Feng Chien
CIKM1
2005 Annotating Text Segments in Documents for Search
abstract
It has been shown that annotating prominent text patterns contained in documents with appropriate types may benefit many applications. Most conventional tools for automatic text annotation extract named entities from texts and annotate them with information about persons, locations, dates and so on. However, this kind of entity type information is often short in length and is mostly limited to a small set of broader categories. In this paper, we try to remedy this problem by presenting an approach to extract global evidences from documents for improved named entity recognition. We also propose an unsupervised, generalized classification approach that collects training data from the Web automatically and classifies text patterns into more refined categories. Experimental results show the feasibility of the proposed approaches for search on the data of the NTCIR-2 information retrieval task.
Pu-Jen Cheng, Hsin-Chen Chiao, Yi-Cheng Pan, Lee-Feng Chien
Web Intelligence1
2004 Creating Multilingual Translation Lexicons with Regional Variations Using Web Corpora
abstract
The purpose of this paper is to automatically create multilingual translation lexicons with regional variations. We propose a transitive translation approach to determine translation variations across languages that have insufficient corpora for translation via the mining of bilingual search-result pages and clues of geographic information obtained from Web search engines. The experimental results have shown the feasibility of the proposed approach in efficiently generating translation equivalents of various terms not covered by general translation dictionaries. It also revealed that the created translation lexicons can reflect different cultural aspects across regions such as Taiwan, Hong Kong and mainland China.
Pu-Jen Cheng, Wen-Hsiang Lu, Jei-Wen Teng, Lee-Feng Chien
ACL1
2004 Translating unknown queries with web corpora for cross-language information retrieval
abstract
It is crucial for cross-language information retrieval (CLIR) systems to deal with the translation of unknown queries due to that real queries might be short. The purpose of this paper is to investigate the feasibility of exploiting the Web as the corpus source to translate unknown queries for CLIR. We propose an online translation approach to determine effective translations for unknown query terms via mining of bilingual search-result pages obtained from Web search engines. This approach can alleviate the problem of the lack of large bilingual corpora, translate many unknown query terms, provide flexible query specifications, and extract semantically-close translations to benefit CLIR tasks -- especially for cross-language Web search.
Pu-Jen Cheng, Jei-Wen Teng, Ruey-Cheng Chen, Jenq-Haur Wang, Wen-Hsiang Lu, Lee-Feng Chien
SIGIR1
2003 Auto-generation of topic hierarchies for web images from users' perspectives
abstract
In this paper, we propose an approach to automatically generating a Yahoo!-like topic hierarchy for organizing Web images from users' perspectives. Relatively little effort has been devoted towards providing such a taxonomy simultaneously considering users' image requests for semantic and visual information. Based on the characteristic that a Web-image query may be refined by various attributes, the proposed approach hierarchically groups similar queries from search engine logs into topic classes at different semantic levels. The generated topic hierarchy has the advantages of organizing image data from users' perspectives for browsing, searching, annotation and users' needs analysis.A series of experiments have been conducted on real-world image search engine logs. Experimental results show that the proposed approach is feasible to generate topic hierarchies for Web images. Moreover, the generated hierarchy has been successfully applied to analysis of users' search interests, which have more focuses on some specific domains when compared with document requests.
Pu-Jen Cheng, Lee-Feng Chien
CIKM1