EDBT 2026 Demo / reviewers in the wild / expert
Lee-Feng Chien
dblp:36/1846
· DBLP profile ↗
62ranked-venue papers
11as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 3 first-authorDatabases, data management, data science and information retrieval · 27 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
12 papers |
Information retrieval · 79% Data mining · 10% Knowledge graphs · 6% | |
| Artificial intelligence
5 papers |
Machine translation · 49% Information extraction and text analysis · 42% Speech recognition and synthesis · 4% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 29 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
cross-language information retrieval |
0.2 | 5 | 2004 | Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004 Translating unknown queries with web corpora for cross-language information retrieval · SIGIR 2004 Anchor Text Mining for Translation Extraction of Query Terms · SIGIR 2001 |
Natural language and speech › Machine translation
named entity translation |
0.1 | 1 | 2007 | Named Entity Translation with Web Mining and Transliteration · IJCAI 2007 |
Multimedia analysis and retrieval › near-duplicate detection
image copy detection |
0.1 | 1 | 2007 | A New Approach to Image Copy Detection Based on Extended Feature Sets · IEEE Trans. Image Process. 2007 |
Information retrieval › cross-language information retrieval
query translation |
0.1 | 2 | 2004 | Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004 Automatic Acquisition of Phrasal Knowledge for English-Chinese Bilingual Information Retrieval · SIGIR 1998 |
Knowledge graphs
taxonomy construction |
0.1 | 1 | 2005 | Taxonomy generation for text segments: A practical web-based approach · ACM Trans. Inf. Syst. 2005 |
Data mining › text mining › text classification
hierarchical text classification |
0.0 | 1 | 2004 | Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004 |
Information retrieval
multilingual dictionary construction |
0.0 | 1 | 2004 | Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004 |
Data mining › text mining
text classification |
0.0 | 1 | 2004 | Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004 |
Information retrieval
web corpus |
0.0 | 1 | 2004 | Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004 |
Information retrieval › query understanding
query analysis |
0.0 | 1 | 2002 | Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering Approach · ICDM 2002 |
Information retrieval › query understanding
query clustering |
0.0 | 1 | 2002 | Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering Approach · ICDM 2002 |
Information retrieval
query processing |
0.0 | 2 | 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000 Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995 |
Information retrieval
web search |
0.0 | 2 | 2004 | Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004 Translating unknown queries with web corpora for cross-language information retrieval · SIGIR 2004 |
Information retrieval
interactive information retrieval |
0.0 | 1 | 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000 |
Information retrieval
query log analysis |
0.0 | 1 | 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000 |
Information retrieval › text analysis
thesaurus construction |
0.0 | 1 | 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000 |
Information retrieval › cross-language information retrieval
chinese text retrieval |
0.0 | 2 | 1997 | Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995 PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997 |
Information retrieval
indexing |
0.0 | 1 | 1997 | PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997 |
Information retrieval › text analysis
keyword extraction |
0.0 | 1 | 1997 | PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997 |
Web and social media mining
web mining |
0.0 | 1 | 2004 | Mining the Web for Generating Thematic Metadata from Textual Data · ICDE 2004 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 1 | 1993 | A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993 |
Natural language and speech › Language models and text generation
language modeling |
0.0 | 1 | 1993 | A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993 |
Automata and formal languages
grammar formalisms |
0.0 | 1 | 1993 | A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993 |
Automata and formal languages › grammar formalisms
unification grammars |
0.0 | 1 | 1993 | A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993 |
Information retrieval
query understanding |
0.0 | 1 | 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
lattice parsing |
0.0 | 1 | 1991 | A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991 |
Compilers and program optimization › parsing
chart parsing |
0.0 | 1 | 1991 | A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991 |
Programming languages and type systems › grammar formalisms
unification grammar |
0.0 | 1 | 1991 | A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991 |
Information retrieval › query formulation
natural language query |
0.0 | 1 | 1995 | Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995 |
Methods — techniques the papers use, named apart from their topics
anchor text mining · 0.1web search · 0.1transitive translation · 0.1text categorization · 0.1geographic information mining · 0.1feature extraction · 0.1bilingual search-result page mining · 0.1web mining · 0.1virtual prior attacks · 0.1transliteration · 0.1classifier learning · 0.1web-based enrichment · 0.1hierarchical clustering · 0.1web corpus mining · 0.0information extraction · 0.0competitive linking · 0.0bilingual lexicon extraction · 0.0best-first chart parsing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Managing and mining multilingual documents: Introduction to the special topic issue of information processing management
Christopher C. Yang, Chih-Ping Wei, Lee-Feng Chien |
Inf. Process. Manag. | 3 |
| 2008 | A Framework for Handling Spatiotemporal Variations in Video Copy DetectionabstractAn effective video copy detection framework should be robust against spatial and temporal variations, e.g., changes in brightness and speed. To this end, a content-based approach for video copy detection is proposed. We define the problem as a partial matching problem in a probabilistic model and transform it into a shortest-path problem in a matching graph. To reduce the computation costs of the proposed framework, we introduce some methods that rapidly select key frames and candidate segments from a large amount of video data. The experiment results show that the proposed approach not only handles spatial and temporal variations well, but it also reduces the computation costs substantially. Chih-Yi Chiu, Chu-Song Chen, Lee-Feng Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Named Entity Translation with Web Mining and Transliteration
Long Jiang, Ming Zhou 0001, Lee-Feng Chien, Cheng Niu |
IJCAI | 3 |
| 2007 | Web-based text classification in the absence of manually labeled training documentsabstractAbstract Most text classification techniques assume that manually labeled documents (corpora) can be easily obtained while learning text classifiers. However, labeled training documents are sometimes unavailable or inadequate even if they are available. The goal of this article is to present a self‐learned approach to extract high‐quality training documents from the Web when the required manually labeled documents are unavailable or of poor quality. To learn a text classifier automatically, we need only a set of user‐defined categories and some highly related keywords. Extensive experiments are conducted to evaluate the performance of the proposed approach using the test set from the Reuters‐21578 news data set. The experiments show that very promising results can be achieved only by using automatically extracted documents from the Web. Chen-Ming Hung, Lee-Feng Chien |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | A New Approach to Image Copy Detection Based on Extended Feature SetsabstractConventional image copy detection research concentrates on finding features that are robust enough to resist various kinds of image attacks. However, finding a globally effective fealure is difficult and, in many cases, domain dependent. Instead of imply extracting features from copyrighted images directly, we propose a new framework called the extended feature set for detecting copies of images. In our approach, virtual prior attacks are applied to copyrighted images to generate novel features, which serve as training data. The copy-detection problem can be solved by learning classifiers from the training data, thus, generated. Our approach can be integrated into existing copy detectors to further improve their performance. Experiment results demonstrate that the proposed approach can substantially enhance the accuracy of copy detection. Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen |
IEEE Trans. Image Process. | 3 |
| 2006 | Query taxonomy generation for web searchabstractWe propose an approach that organizes the search-result clusters into a hierarchical structure, called a query taxonomy, from the user's perspective. The proposed approach is based on an unsupervised classification method, which uses the dynamic Web as the training corpus. With query taxonomy, users can browse relevant Web documents more conveniently and comprehensibly. Our experimental results verify the feasibility and the effectiveness of the proposed approach to query taxonomy generation in Web search. Pu-Jen Cheng, Ching-Hsiang Tsai, Chen-Ming Hung, Lee-Feng Chien |
CIKM | 4 |
| 2006 | Image Copy Detection via Grouping in Feature Space Based on Virtual Prior AttacksabstractIn the past, many researches on image copy detection focused on finding a feature that is robust enough for various kinds of image attacks. But it is difficult to find a globally effective feature that is appropriate for many situations. In this paper, we introduce a classification framework to this problem, instead of solving the feature-selection problem. In our approach, novel features are generated by applying virtual prior attacks to copyrighted images, and the copy-detection problem is converted to a classification one that is more robust to solve. Our approach can combine existing image copy detectors and further raise their performances. Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen |
ICIP | 3 |
| 2006 | Image Content Clustering and Summarization for Photo CollectionsabstractRapid growth of digital photography in recent years spurred the need of photo management tools. In this study, we propose an automatic organization framework for photo collections based on image content, so that a novel browsing experience is provided for users. For each photograph, human faces, together with corresponding clothes and nearby regions are located. We extract color histograms of these regions as the image content feature. Then a similarity matrix of a photo collection is generated according to temporal and content features of those photographs. We perform hierarchical clustering based on this matrix, and extract duplicate subjects of a cluster by introducing the contrast context histogram (CCH) technique. The experimental results show that the developed framework provides a promising result for photo management Cheng-Hung Li, Chih-Yi Chiu, Chun-Rong Huang, Chu-Song Chen, Lee-Feng Chien |
ICME | 5 |
| 2006 | Exploiting the Web as the multilingual corpus for unknown query translationabstractAbstract Users' cross‐lingual queries to a digital library system might be short and the query terms may not be included in a common translation dictionary (unknown terms). In this article, the authors investigate the feasibility of exploiting the Web as the multilingual corpus source to translate unknown query terms for cross‐language information retrieval in digital libraries. They propose a Web‐based term translation approach to determine effective translations for unknown query terms by mining bilingual search‐result pages obtained from a real Web search engine. This approach can enhance the construction of a domain‐specific bilingual lexicon and bring multilingual support to a digital library that only has monolingual document collections. Very promising results have been obtained in generating effective translation equivalents for many unknown terms, including proper nouns, technical terms, and Web query terms, and in assisting bilingual lexicon construction for a real digital library system. Jenq-Haur Wang, Jei-Wen Teng, Wen-Hsiang Lu, Lee-Feng Chien |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2005 | Annotating Text Segments in Documents for SearchabstractIt has been shown that annotating prominent text patterns contained in documents with appropriate types may benefit many applications. Most conventional tools for automatic text annotation extract named entities from texts and annotate them with information about persons, locations, dates and so on. However, this kind of entity type information is often short in length and is mostly limited to a small set of broader categories. In this paper, we try to remedy this problem by presenting an approach to extract global evidences from documents for improved named entity recognition. We also propose an unsupervised, generalized classification approach that collects training data from the Web automatically and classifies text patterns into more refined categories. Experimental results show the feasibility of the proposed approaches for search on the data of the NTCIR-2 information retrieval task. Pu-Jen Cheng, Hsin-Chen Chiao, Yi-Cheng Pan, Lee-Feng Chien |
Web Intelligence | 4 |
| 2005 | Automatic Training Corpora Acquisition through Web MiningabstractText classification is a task having been extensively studied for decades. However, most previous work pre-assumes the existence of explicitly labeled corpora. In this study, we focus on the issue of automatic corpora acquisition. We propose a Web-based mining approach to collect necessary corpora, which can be greatly useful to both common users and system designers. Moreover, the proposed technique can also be incorporated with existing classification techniques to further boost classifier performance. It has been shown that the concept of the class can be captured by the class name and its associated terms (Huang et al., 2004). In this work, we aim at analyzing Web-retrieved documents to discover the associated terms, which are further utilized to collect more training corpora. Working iteratively, the proposed approach can acquire training corpora of high quality. We give empirical evidence that the classifiers thus created have promising accuracy. In sum, the convenience and efficiency of the proposed approach, along with the new perspective on the issue of corpora acquisition, are the primary contributions of this work. Kuan-Ming Lin, Lee-Feng Chien |
Web Intelligence | 3 |
| 2005 | Taxonomy generation for text segments: A practical web-based approachabstractIt is crucial in many information systems to organize short text segments, such as keywords in documents and queries from users, into a well-formed taxonomy. In this article, we address the problem of taxonomy generation for diverse text segments with a general and practical approach that uses the Web as an additional knowledge source. Unlike long documents, short text segments typically do not contain enough information to extract reliable features. This work investigates the possibilities of using highly ranked search-result snippets to enrich the representation of text segments. A hierarchical clustering algorithm is then designed for creating the hierarchical topic structure of text segments. Text segments with close concepts can be grouped together in a cluster, and relevant clusters linked at the same or near levels. Different from traditional clustering algorithms, which tend to produce cluster hierarchies with a very unnatural shape, the algorithm tries to produce a more natural and comprehensive tree hierarchy. Extensive experiments were conducted on different domains of text segments, including subject terms, people names, paper titles, and natural language questions. The obtained experimental results have shown the potential of the proposed approach, which provides a basis for the in-depth analysis of text segments on a larger scale and is believed able to benefit many information systems. Shui-Lung Chuang, Lee-Feng Chien |
ACM Trans. Inf. Syst. | 2 |
| 2004 | Creating Multilingual Translation Lexicons with Regional Variations Using Web CorporaabstractThe purpose of this paper is to automatically create multilingual translation lexicons with regional variations. We propose a transitive translation approach to determine translation variations across languages that have insufficient corpora for translation via the mining of bilingual search-result pages and clues of geographic information obtained from Web search engines. The experimental results have shown the feasibility of the proposed approach in efficiently generating translation equivalents of various terms not covered by general translation dictionaries. It also revealed that the created translation lexicons can reflect different cultural aspects across regions such as Taiwan, Hong Kong and mainland China. Pu-Jen Cheng, Wen-Hsiang Lu, Jei-Wen Teng, Lee-Feng Chien |
ACL | 4 |
| 2004 | A practical web-based approach to generating topic hierarchy for text segmentsabstractIt is crucial in many information systems to organize short text segments, such as keywords in documents and queries from users, into a well-formed topic hierarchy. In this paper, we address the problem of generating topic hierarchies for diverse text segments with a general and practical approach that uses the Web as an additional knowledge source. Unlike long documents, short text segments typically do not contain enough information to extract reliable features. This work investigates the possibilities of using highly ranked search-result snippets to enrich the representation of text segments. A hierarchical clustering algorithm is then applied to create the hierarchical topic structure of text segments. Different from traditional clustering algorithms, which tend to produce cluster hierarchies with a very unnatural shape, the approach tries to produce a more natural and comprehensive hierarchy. Extensive experiments were conducted on different domains of text segments. The obtained results have shown the potential of the proposed approach, which is believed able to benefit many information systems. Shui-Lung Chuang, Lee-Feng Chien |
CIKM | 2 |
| 2004 | Mining the Web for Generating Thematic Metadata from Textual DataabstractConventional tools for automatic metadata creation mostly extract named entities or patterns from texts and annotate them with information about persons, locations, dates, and so on. However, this kind of entity type information is often too primitive for more advanced intelligent applications such as concept-based search. Here, we try to generate semantically-deep metadata with limited human intervention. The main idea behind our approach is to use Web mining and categorization techniques to create thematic metadata. The proposed approach, comprises of three computational modules: feature extraction, HCQF (hier-concept query formulation) and text instance categorization. The feature extraction module sends the name of text instances to Web search engines, and the returned highly-ranked search-result pages are used to describe them. Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien |
ICDE | 3 |
| 2004 | Categorizing Unknown Text Segments for Information Extraction Using a Search Result Mining Approach
Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien |
IJCNLP | 3 |
| 2004 | Generating Concept Hierarchies from Text for Intelligence Analysis
Jenq-Haur Wang, Chien-Chung Huang 0004, Jei-Wen Teng, Lee-Feng Chien |
ISI | 4 |
| 2004 | Translating unknown queries with web corpora for cross-language information retrievalabstractIt is crucial for cross-language information retrieval (CLIR) systems to deal with the translation of unknown queries due to that real queries might be short. The purpose of this paper is to investigate the feasibility of exploiting the Web as the corpus source to translate unknown queries for CLIR. We propose an online translation approach to determine effective translations for unknown query terms via mining of bilingual search-result pages obtained from Web search engines. This approach can alleviate the problem of the lack of large bilingual corpora, translate many unknown query terms, provide flexible query specifications, and extract semantically-close translations to benefit CLIR tasks -- especially for cross-language Web search. Pu-Jen Cheng, Jei-Wen Teng, Ruey-Cheng Chen, Jenq-Haur Wang, Wen-Hsiang Lu, Lee-Feng Chien |
SIGIR | 6 |
| 2004 | Liveclassifier: creating hierarchical text classifiers through web corporaabstractMany Web information services utilize techniques of information extraction(IE) to collect important facts from the Web. To create more advanced services, one possible method is to discover thematic information from the collected facts through text classification. However, most conventional text classification techniques rely on manual-labelled corpora and are thus ill-suited to cooperate with Web information services with open domains. In this work, we present a system named LiveClassifier that can automatically train classifiersthrough Web corpora based on user-defined topic hierarchies. Due to its flexibility and convenience, LiveClassifier can be easily adapted for various purposes. New Web information services can be created to fully exploit it; human users can use it to create classifiers for their personal applications. The effectiveness of classifiers created by LiveClassifier is well supportedby empirical evidence. Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien |
WWW | 3 |
| 2004 | Using a web-based categorization approach to generate thematic metadata from textsabstractConventional tools for automatic metadata creation mostly extract named entities or text segments from texts and annotate them with information about persons, locations, dates, and so on. However, this kind of entity type information is often insufficient for machines to understand the facts contained in the texts, thus precluding the possibility of implementing more advanced, intelligent applications, such as concept-based search. In this work, we try to create more refined thematic metadata inherent in texts. Based on Web resource mining, our approach acquires training corpora necessary to describe both the thematic categories and the metadata extracted from the texts. The approach then finds the corresponding relationships among them by means of categorization and thus generates thematic metadata for the textual data. Experimental results confirm the potential and wide adaptability of our approach. Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2004 | Anchor text mining for translation of Web queries: A transitive translation approachabstractTo discover translation knowledge in diverse data resources on the Web, this article proposes an effective approach to finding translation equivalents of query terms and constructing multilingual lexicons through the mining of Web anchor texts and link structures. Although Web anchor texts are wide-scoped hypertext resources, not every particular pair of languages contains sufficient anchor texts for effective extraction of translations for Web queries. For more generalized applications, the approach is designed based on a transitive translation model. The translation equivalents of a query term can be extracted via its translation in an intermediate language. To reduce interference from translation errors, the approach further integrates a competitive linking algorithm into the process of determining the most probable translation. A series of experiments has been conducted, including performance tests on term translation extraction, cross-language information retrieval, and translation suggestions for practical Web search services, respectively. The obtained experimental results have shown that the proposed approach is effective in extracting translations of unknown queries, is easy to combine with the probabilistic retrieval model to improve the cross-language retrieval performance, and is very useful when the considered language pairs lack a sufficient number of anchor texts. Based on the approach, an experimental system called LiveTrans has been developed for English--Chinese cross-language Web search. Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee |
ACM Trans. Inf. Syst. | 2 |
| 2003 | Auto-generation of topic hierarchies for web images from users' perspectivesabstractIn this paper, we propose an approach to automatically generating a Yahoo!-like topic hierarchy for organizing Web images from users' perspectives. Relatively little effort has been devoted towards providing such a taxonomy simultaneously considering users' image requests for semantic and visual information. Based on the characteristic that a Web-image query may be refined by various attributes, the proposed approach hierarchically groups similar queries from search engine logs into topic classes at different semantic levels. The generated topic hierarchy has the advantages of organizing image data from users' perspectives for browsing, searching, annotation and users' needs analysis.A series of experiments have been conducted on real-world image search engine logs. Experimental results show that the proposed approach is feasible to generate topic hierarchies for Web images. Moreover, the generated hierarchy has been successfully applied to analysis of users' search interests, which have more focuses on some specific domains when compared with document requests. Pu-Jen Cheng, Lee-Feng Chien |
CIKM | 2 |
| 2003 | Enriching Web taxonomies through subject categorization of query terms from search engine logs
Shui-Lung Chuang, Lee-Feng Chien |
Decis. Support Syst. | 2 |
| 2003 | Relevant term suggestion in interactive web search based on contextual information in query session logsabstractAbstract This paper proposes an effective term suggestion approach to interactive Web search. Conventional approaches to making term suggestions involve extracting co‐occurring keyterms from highly ranked retrieved documents. Such approaches must deal with term extraction difficulties and interference from irrelevant documents, and, more importantly, have difficulty extracting terms that are conceptually related but do not frequently co‐occur in documents. In this paper, we present a new, effective log‐based approach to relevant term extraction and term suggestion. Using this approach, the relevant terms suggested for a user query are those that co‐occur in similar query sessions from search engine logs, rather than in the retrieved documents. In addition, the suggested terms in each interactive search step can be organized according to its relevance to the entire query session, rather than to the most recent single query as in conventional approaches. The proposed approach was tested using a proxy server log containing about two million query transactions submitted to search engines in Taiwan. The obtained experimental results show that the proposed approach can provide organized and highly relevant terms, and can exploit the contextual information in a user's query session to make more effective suggestions. Chien-Kang Huang, Lee-Feng Chien, Yen-Jen Oyang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2002 | A Transitive Model for Extracting Translation Equivalents of Web Queries through Anchor Text Mining
Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee |
COLING | 2 |
| 2002 | Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering ApproachabstractMost previous work on automatic query clustering generated a flat, un-nested partition of query terms. In this work, we discuss the organization of query terms into a hierarchical structure and construct a query taxonomy in an automatic way. The proposed approach is designed based on a hierarchical agglomerative clustering algorithm to hierarchically group similar queries and generate cluster hierarchies using a novel cluster partition technique. The search processes of real-world search engines are combined to obtain highly ranked Web documents as the feature source for each query term. Preliminary experiments show that the proposed approach is effective for obtaining thesaurus information for query terms, and is also feasible for constructing a query taxonomy which provides a basis for in-depth analysis of users' search interests and domain-specific vocabulary on a larger scale. Shui-Lung Chuang, Lee-Feng Chien |
ICDM | 2 |
| 2002 | Incremental Extraction of Keyterms for Classifying Multilingual Documents in the Web
Lee-Feng Chien, Chien-Kang Huang, Hsin-Chen Chiao, Shih-Jui Lin |
PAKDD | 1 |
| 2002 | Translation of web queries using anchor text miningabstractThis article presents an approach to automatically extracting translations of Web query terms through mining of Web anchor texts and link structures. One of the existing difficulties in cross-language information retrieval (CLIR) and Web search is the lack of appropriate translations of new terminology and proper names. The proposed approach successfully exploits the anchor-text resources and reduces the existing difficulties of query term translation. Many query terms that cannot be obtained in general-purpose translation dictionaries are, therefore, extracted. Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2001 | Anchor Text Mining for Translation of Web QueriesabstractThe paper presents an approach to automatically extracting translations of Web query terms through mining of Web anchor texts and link structures. One of the existing difficulties in cross-language information retrieval (CLIR) and Web search is the lack of the appropriate translations of new terminology and proper names. Such a difficult problem can be effectively alleviated by our proposed approach, and the resource of anchor texts in the Web is proven a valuable corpus for this kind of term translation. Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee |
ICDM | 2 |
| 2001 | Anchor Text Mining for Translation Extraction of Query TermsabstractThis paper presents an approach to automatically extracting the bilingual translations of many Web query terms through mining the Web anchor texts. Some preliminary experiments are conducted on using 109,416 Web pages containing both Chinese and English anchor texts in their in-links to extract Chinese translations of 200 English queries selected from popular query terms in Taiwan. It is found that the effective translations of 75% of the popular query terms can be extracted, in which 87.2% cannot be obtained in common translation dictionaries. Wen-Hsiang Lu, Hsi-Jian Lee, Lee-Feng Chien |
SIGIR | 3 |
| 2001 | A Contextual Term Suggestion Mechanism for Interactive Web Search
Chien-Kang Huang, Yen-Jen Oyang, Lee-Feng Chien |
Web Intelligence | 3 |
| 2000 | Live thesaurus construction for interactive voice-based web searchabstractSince Web users ’ queries are often too short, an accurate and interactive speech interface is believed very helpful, especially for WAP-based Web search engines. To provide high accurate speech recognition and effective interactive search, a rigid and live web thesaurus that contains users ' search terms plus a set of relations between their associated terms is highly in demand. The purpose of this paper is intended to present a log-based approach for live thesaurus construction. Based on the live thesaurus, certain kinds of users ' information behaviors could be characterized and a more effective voice-based search engine could be developed. Shui-Lung Chuang, Hsiao-Tieh Pu, Wen-Hsiang Lu, Lee-Feng Chien |
INTERSPEECH | 4 |
| 2000 | Live Lexicons and Dynamic Corpora Adapted to the Network Resources for Chinese Spoken Language Processing Applications in an Internet Era
Lin-Shan Lee, Lee-Feng Chien |
LREC | 2 |
| 2000 | Auto-construction of a live thesaurus from search term logs for interactive Web searchabstractThe purpose of this paper is to present an on-going research that is intended to construct a live thesaurus directly from search term logs of real-world search engines. Such a thesaurus designed can contain representative search terms, their frequency in use, the corresponding subject categories, the associated and relevant terms, and the hot visiting Web sites/pages the search terms may reach. Shui-Lung Chuang, Hsiao-Tieh Pu, Wen-Hsiang Lu, Lee-Feng Chien |
SIGIR | 4 |
| 2000 | A spoken-access approach for chinese text and speech information retrievalabstractThis paper presents an efficient spoken-access approach for both Chinese text and Mandarin speech information retrieval. The proposed approach is developed not only to deal with the retrieval of spoken documents, but also to improve the capability of human-computer interaction via voice input for information-retrieval systems. Based on utilization of the monosyllabic structure of the Chinese language, the proposed approach can tolerate speech recognition errors by performing speech query recognition and approximate information retrieval at the syllable-level. Furthermore, with the help of automatic term suggestion and relevance feedback techniques, the proposed approach is robust in enabling users using voice input to interact with IR systems at each stage of the retrieval process. Extensive experiments show that the proposed approach can improve the effectiveness of information retrieval via speech interaction. The encouraging results suggest that a Mandarin speech interface for information retrieval and digital library systems can, therefore, be developed. Lee-Feng Chien, Hsin-Min Wang, Bo-Ren Bai, Sun-Chien Lin |
J. Am. Soc. Inf. Sci. | 1 |
| 1999 | An OODBMS-IRS Integration Based on a Statistical Corpus Extraction Method for Document Management
Chung-Hong Lee, Lee-Feng Chien |
DEXA | 2 |
| 1999 | PAT-tree-based adaptive keyphrase extraction for intelligent Chinese information retrieval
Lee-Feng Chien |
Inf. Process. Manag. | 1 |
| 1998 | A Novel Integration of OODBMS and Information Retrieval Techniques for a Document Repository
Chung-Hong Lee, Lee-Feng Chien |
DEXA | 2 |
| 1998 | Statistics-based segment pattern lexicon-a new direction for Chinese language modelingabstractThis paper presents a new direction for Chinese language modeling based on a different concept of the lexicon. Because every Chinese character has its own meaning and there are no "blanks" in Chinese sentences serving as word boundaries, also because the wording structure in the Chinese language is extremely flexible, the "words" in Chinese are actually not well defined, and there does not exist a commonly accepted lexicon. This makes language modeling very sophisticated in the Chinese language, and the "out of vocabulary (OOV)" problem specially serious. A new concept for the lexicon is thus proposed. The elements of this lexicon can be words or any other "segment patterns". They should be extracted from the training corpus by statistical approaches with a goal to minimize the overall perplexity. The language models can then be developed based on this new lexicon. Very encouraging experimental results have been obtained. Kae-Cherng Yang, Tai-Hsuan Ho, Lee-Feng Chien, Lin-Shan Lee |
ICASSP | 3 |
| 1998 | A*-admissible key-phrase spotting with sub-syllable level utterance verificationabstractIn this paper, we propose an A*-admissible key-phrase spotting framework, which needs little domain knowledge and is capable of extracting salient key-phrase fragments from an input utterance in real-time. There are two key features in our approach. Firstly, the acoustic models and the search framework are specially designed such that very high degree vocabulary flexibility can be achieved for any desired application tasks. Secondly, the search framework uses an efficient two-pass A* search to generate N-best key-phrase candidates and then several sub-syllable level verification functions are properly weighted and used to further improve the recognition accuracy. Experimental results show that the A*-admissible key-phrase spotting with sub-word level utterance method outperforms the baseline methods used in common approaches. 1. INTRODUCTION In recent years, various spoken dialog systems have been widely investigated for the fast growing demand for real-world applications. It is diff... Berlin Chen, Hsin-Min Wang, Lee-Feng Chien, Lin-Shan Lee |
ICSLP | 3 |
| 1998 | Automatic Acquisition of Phrasal Knowledge for English-Chinese Bilingual Information RetrievalabstractNo abstract available. Ming-Jer Lee, Lee-Feng Chien |
SIGIR | 2 |
| 1997 | Internet Chinese information retrieval using unconstrained Mandarin speech queries based on a client-server architecture and a PAT-tree-based language modelabstractIn order to pursue high performance of Chinese information access on the Internet, this paper presents an attractive approach with a successful integration of efficient speech recognition and information retrieval techniques. A working system based on the proposed approach for speech retrieval of real-time Chinese netnews services has been implemented and tested. Very exciting performance has been achieved. Lee-Feng Chien, Sung-Chien Lin, Jenn-Chau Hong, Ming-Chiuan Chen, Hsin-Min Wang, Jia-Lin Shen, Keh-Jiann Chen, Lin-Shan Lee |
ICASSP | 1 |
| 1997 | Syllable-based relevance feedback techniques for Mandarin voice record retrieval using speech queriesabstractIn order to solve the problem with the new environment of fast growth of audio resources on the Internet, we have presented a syllable-based approach which is capable of retrieving Mandarin voice records using queries of unconstrained speech. However, the performance achieved by this previously proposed approach is still not satisfactory, and one of the reason is that very often the information provided by the speech query for the request subject may not be sufficient. We present approaches based the relevance feedback technique to improving the performance of the previous research. The proposed approaches include a relevance measure adjustment scheme using a relevance table for the voice database, a query expansion scheme to generate a new query including the feedback information, and a combination of these two schemes. Extensive preliminary experiments were performed and demonstrated. Lin-Shan Lee, Bo-Ren Bai, Lee-Feng Chien |
ICASSP | 3 |
| 1997 | Intelligent retrieval of very large Chinese dictionaries with speech queriesabstractTo retrieve a Chinese word from a Chinese dictionary, it needs the user to know exactly the first character of the desired word. Because there is more than 10,000 Chinese characters, this makes the Chinese dictionary relatively difficult to be used. To reduce the problem, this paper presents intelligent retrieval techniques for very large Chinese dictionaries with speech queries. The proposed techniques properly integrate the technologies of Mandarin speech recognition and Chinese information retrieval with a syllable-based approach utilizing the mono-syllabic structure of the language. Moreover, it is very nice to provide the function of retrieving all relevant word entries from the dictionaries using speech queries describing “general concepts” of the desired words. To achieve the challenging function, the techniques of relevance feedback are also included. Based on these techniques, a retrieval system was implemented successfully on a Pentium PC for a very large Chinese dictionary which includes 160,000 word entries and the total length of the lexical information under the word entries exceeds 20,000,000 words. Sung-Chien Lin, Lee-Feng Chien, Ming-Chiuan Chen, Lin-Shan Lee, Keh-Jiann Chen |
EUROSPEECH | 2 |
| 1997 | Chinese language model adaptation based on document classification and multiple domain-specific language models
Sung-Chien Lin, Chi-Lung Tsai, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
EUROSPEECH | 3 |
| 1997 | PAT-tree-based Keyword Extraction for Chinese Information Retrievalabstracturgent need to promote Chinese in this paper we will raise the significance of keyword extraction using a new PAT-treebased approach, which is efficient in automatic keyword extraction from a set of relevant Chinese documents.This approach has been successfully applied in several IR researches, such as document classification, book indexing and relevance feedback.Many Chinese language processing applications therefore step ahead from character level to word/phrase level, Lee-Feng Chien |
SIGIR | 1 |
| 1996 | An efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a clustered language modelabstractThis paper presents an accurate and efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a specially-designed clustered language model. To reduce the problems resulted from the complexity of unconstrained speech-input queries for retrieval, the system is completely syllable-based in both speech recognition and database retrieval by properly utilizing the mono-syllabic structure of Chinese language. In addition, it partitions the records in the database into clusters and trains the clustered language model using the clustering results. The proposed clustered language model with its augmented search algorithm are very useful to improve accuracy and speed of the speech retrieval system. In the preliminary tests using an experimental database with about 30,000 bibliographical records, it was found that the present system can accept unconstrained speech-input queries and achieve very good performance. Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
ICASSP | 2 |
| 1996 | Very-large-vocabulary Mandarin voice message file retrieval using speech queriesabstractIn order to solve the problem with the new environment of fast growth of audio resources on the Intemet, this paper presents a new approach which is capable of retrieving Mandarin voice message files using queries of unconstrained speech.By properly utilizing the monosyllabic structure of the Chinese language, the proposed approach perfoms the statistical similarity estimation between the speech queries and the voice message files, and executes the complete matching process directly at the phonetic level using syllable-based statistical information.Based on this approach, some experiments are tested and encouraging results are demonstrated. Bo-Ren Bai, Lee-Feng Chien, Lin-Shan Lee |
ICSLP | 2 |
| 1996 | Speaker intention modeling for large vocabulary Mandarin spoken dialoguesabstractThis paper presents a statistical speaker intention modeling approach of speech act types (SAT's)[l] prediction for large vocabulary Mandarin spoken dialogues.A SAT is an abstraction of speaker's intention in terms of the type of action thax the speaker intends by the utterance.With this approach, spoken dialogue systems can be constructed to predict speaker's intention and make a proper action in advance. Yen-Ju Yang, Lee-Feng Chien, Lin-Shan Lee |
ICSLP | 2 |
| 1995 | Golden Mandarin (III)-a user-adaptive prosodic-segment-based Mandarin dictation machine for Chinese language with very large vocabularyabstractThis paper presents a prototype prosodic-segment-based Mandarin dictation machine for the Chinese language with very large vocabulary. It accepts utterances continuous within a prosodic segment which is composed of one or a few word(s). It also possesses various on-line learning capabilities for fast adaptation to a new user in acoustic, lexical and linguistic levels. The overall system is implemented on an IBM/PC with an additional DSP card including a Motorola DSP 96002 chip. The word accuracy can achieve nearly 90% for a new user after he produces about 10 minutes of speech to train the system, and the accuracy can be further improved with the on-line learning functions. Ren-Yuan Lyu, Lee-Feng Chien, Shiao-Hong Hwang, Hung-Yun Hsieh, Rung-Chiuan Yang, Bo-Ren Bai, Jia-Chi Weng, Yen-Ju Yang, Shi-Wei Lin, Keh-Jiann Chen, Chiu-yu Tseng, Lin-Shan Lee |
ICASSP | 2 |
| 1995 | Fast and accurate continuous speech recognition for Chinese language with very large vocabulary
Tai-Hsuan Ho, Hsin-Min Wang, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
EUROSPEECH | 3 |
| 1995 | A syllable-based very-large-vocabulary voice retrieval system for Chinese databases with textual attributes
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
EUROSPEECH | 2 |
| 1995 | Unconstrained speech retrieval for Chinese document databases with very large vocabulary and unlimited domains
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
EUROSPEECH | 2 |
| 1995 | Fast and Quasi-Natural Language Search for Gigabits of Chinese TextsabstractArticle Fast and quasi-natural language search for gigabytes of Chinese texts Share on Author: Lee-Feng Chien Institute of Information Science, Academia Sinica, Taipei, Taiwan, R.O.C. Institute of Information Science, Academia Sinica, Taipei, Taiwan, R.O.C.View Profile Authors Info & Claims SIGIR '95: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrievalJuly 1995 Pages 112–120https://doi.org/10.1145/215206.215345Online:01 July 1995Publication History 26citation508DownloadsMetricsTotal Citations26Total Downloads508Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Lee-Feng Chien |
SIGIR | 1 |
| 1994 | An intelligent and efficient word-class-based Chinese language model for Mandarin speech recognition with very large vocabulary
Yen-Ju Yang, Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
ICSLP | 3 |
| 1993 | Golden Mandarin (II)-an improved single-chip real-time Mandarin dictation machine for Chinese language with very large vocabulary
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen, I-Jung Hung, Ming-Yu Lee, Lee-Feng Chien, Yumin Lee, Ren-Yuan Lyu, Hsin-Min Wang, Yung-Chuan Wu, Tung-Sheng Lin, Hung-Yan Gu, Chi-ping Nee, Chun-Yi Liao, Yeng-Ju Yang, Yuan-Cheng Chang, Rung-Chiung Yang |
ICASSP (2) | 6 |
| 1993 | A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applicationsabstractA language processing model is proposed in which the grammatical approach of unification grammar and the statistical approach of Markov language models are properly integrated in a word lattice chart parsing algorithm with different best-first parsing strategies. This model has been successfully implemented in experiments on Mandarin speech recognition although it is language-independent. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved. A correct rate of 93.8% and 5 s per sentence on an IBM PC/AT, as compared with 73.8% and 25 s using unification grammar alone and 82.2% and 3 s using a Markov language model alone, was achieved. This high performance is due to the effective rejection of noisy word hypothesis interferences; that is, the unification-based grammatical analysis eliminates all illegal combinations, while the Markovian probabilities of constituents combined with the considerations on constituent length indicate the correct direction of processing.> Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 1991 | A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition ApplicationsabstractA language processor is to find out a most promising sentence hypothesis for a given word lattice obtained from acoustic signal recognition. In this paper a new language processor is proposed, in which unification grammar and Markov language model are integrated in a word lattice parsing algorithm based on an augmented chart, and the island-driven parsing concept is combined with various preference-first parsing strategies defined by different construction principles and decision rules. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved. Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
ACL | 1 |
| 1991 | An Efficient Natural Language Processing System Specially Designed for the Chinese Language
Lin-Shan Lee, Lee-Feng Chien, Long Ji Lin, Keh-Jiann Chen |
Comput. Linguistics | 2 |
| 1991 | An augmented chart data structure with efficient word lattice parsing scheme in speech recognition applications
Lee-Feng Chien, Lin-Shan Lee, Keh-Jiann Chen |
Speech Commun. | 1 |
| 1990 | An Augmented Chart Data Structure with Efficient Word Lattice Parsing Scheme In Speech Recognition Applications
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
COLING | 1 |
| 1990 | An augmented chart parsing algorithm integrating unification grammar and Markov language model for continuous speech recognitionabstractAn efficient algorithm is developed to handle the difficulties in parsing noise word lattices (sets of word hypotheses obtained in continuous-speech recognition) which include problems such as word boundary overlapping, homonyms, lexical ambiguities, recognition uncertainty and errors, etc. An augmented chart is proposed, and the algorithms is then derived on this chart. This algorithm properly integrates the global structural synthesis capabilities of the unification grammar and the local relation estimation capabilities of the Markov language model. The parsing algorithm is island driven and best first. In this way, the features of the grammatical and statistical approaches can be combined, and the effects of the two different approaches are reflected in a single algorithm such that the overall selectivity can be appropriately optimized.> Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
ICASSP | 1 |