Lee-Feng Chien

dblp:36/1846 · DBLP profile ↗
← Back
62ranked-venue papers
11as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 3 first-authorDatabases, data management, data science and information retrieval · 27 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Information retrieval · 79% Data mining · 10% Knowledge graphs · 6%
Artificial intelligence
5 papers
Machine translation · 49% Information extraction and text analysis · 42% Speech recognition and synthesis · 4%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 29 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
0.252004
Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004
Translating unknown queries with web corpora for cross-language information retrieval · SIGIR 2004
Anchor Text Mining for Translation Extraction of Query Terms · SIGIR 2001
Natural language and speech › Machine translation
named entity translation
0.112007
Named Entity Translation with Web Mining and Transliteration · IJCAI 2007
Multimedia analysis and retrieval › near-duplicate detection
image copy detection
0.112007
A New Approach to Image Copy Detection Based on Extended Feature Sets · IEEE Trans. Image Process. 2007
Information retrieval › cross-language information retrieval
query translation
0.122004
Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004
Automatic Acquisition of Phrasal Knowledge for English-Chinese Bilingual Information Retrieval · SIGIR 1998
Knowledge graphs
taxonomy construction
0.112005
Taxonomy generation for text segments: A practical web-based approach · ACM Trans. Inf. Syst. 2005
Data mining › text mining › text classification
hierarchical text classification
0.012004
Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004
Information retrieval
multilingual dictionary construction
0.012004
Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004
Data mining › text mining
text classification
0.012004
Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004
Information retrieval
web corpus
0.012004
Liveclassifier: creating hierarchical text classifiers through web corpora · WWW 2004
Information retrieval › query understanding
query analysis
0.012002
Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering Approach · ICDM 2002
Information retrieval › query understanding
query clustering
0.012002
Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering Approach · ICDM 2002
Information retrieval
query processing
0.022000
Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000
Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995
Information retrieval
web search
0.022004
Anchor text mining for translation of Web queries: A transitive translation approach · ACM Trans. Inf. Syst. 2004
Translating unknown queries with web corpora for cross-language information retrieval · SIGIR 2004
Information retrieval
interactive information retrieval
0.012000
Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000
Information retrieval
query log analysis
0.012000
Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000
Information retrieval › text analysis
thesaurus construction
0.012000
Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000
Information retrieval › cross-language information retrieval
chinese text retrieval
0.021997
Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995
PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997
Information retrieval
indexing
0.011997
PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997
Information retrieval › text analysis
keyword extraction
0.011997
PAT-tree-based Keyword Extraction for Chinese Information Retrieval · SIGIR 1997
Web and social media mining
web mining
0.012004
Mining the Web for Generating Thematic Metadata from Textual Data · ICDE 2004
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011993
A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993
Natural language and speech › Language models and text generation
language modeling
0.011993
A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993
Automata and formal languages
grammar formalisms
0.011993
A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993
Automata and formal languages › grammar formalisms
unification grammars
0.011993
A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications · IEEE Trans. Speech Audio Process. 1993
Information retrieval
query understanding
0.012000
Auto-construction of a live thesaurus from search term logs for interactive Web search · SIGIR 2000
Natural language and speech › Information extraction and text analysis › syntactic parsing
lattice parsing
0.011991
A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991
Compilers and program optimization › parsing
chart parsing
0.011991
A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991
Programming languages and type systems › grammar formalisms
unification grammar
0.011991
A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications · ACL 1991
Information retrieval › query formulation
natural language query
0.011995
Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts · SIGIR 1995

Methods — techniques the papers use, named apart from their topics

anchor text mining · 0.1web search · 0.1transitive translation · 0.1text categorization · 0.1geographic information mining · 0.1feature extraction · 0.1bilingual search-result page mining · 0.1web mining · 0.1virtual prior attacks · 0.1transliteration · 0.1classifier learning · 0.1web-based enrichment · 0.1hierarchical clustering · 0.1web corpus mining · 0.0information extraction · 0.0competitive linking · 0.0bilingual lexicon extraction · 0.0best-first chart parsing · 0.0
YearPublicationVenuePosition
2011 Managing and mining multilingual documents: Introduction to the special topic issue of information processing management
Christopher C. Yang, Chih-Ping Wei, Lee-Feng Chien
Inf. Process. Manag.3
2008 A Framework for Handling Spatiotemporal Variations in Video Copy Detection
abstract
An effective video copy detection framework should be robust against spatial and temporal variations, e.g., changes in brightness and speed. To this end, a content-based approach for video copy detection is proposed. We define the problem as a partial matching problem in a probabilistic model and transform it into a shortest-path problem in a matching graph. To reduce the computation costs of the proposed framework, we introduce some methods that rapidly select key frames and candidate segments from a large amount of video data. The experiment results show that the proposed approach not only handles spatial and temporal variations well, but it also reduces the computation costs substantially.
Chih-Yi Chiu, Chu-Song Chen, Lee-Feng Chien
IEEE Trans. Circuits Syst. Video Technol.3
2007 Named Entity Translation with Web Mining and Transliteration
Long Jiang, Ming Zhou 0001, Lee-Feng Chien, Cheng Niu
IJCAI3
2007 Web-based text classification in the absence of manually labeled training documents
abstract
Abstract Most text classification techniques assume that manually labeled documents (corpora) can be easily obtained while learning text classifiers. However, labeled training documents are sometimes unavailable or inadequate even if they are available. The goal of this article is to present a self‐learned approach to extract high‐quality training documents from the Web when the required manually labeled documents are unavailable or of poor quality. To learn a text classifier automatically, we need only a set of user‐defined categories and some highly related keywords. Extensive experiments are conducted to evaluate the performance of the proposed approach using the test set from the Reuters‐21578 news data set. The experiments show that very promising results can be achieved only by using automatically extracted documents from the Web.
Chen-Ming Hung, Lee-Feng Chien
J. Assoc. Inf. Sci. Technol.2
2007 A New Approach to Image Copy Detection Based on Extended Feature Sets
abstract
Conventional image copy detection research concentrates on finding features that are robust enough to resist various kinds of image attacks. However, finding a globally effective fealure is difficult and, in many cases, domain dependent. Instead of imply extracting features from copyrighted images directly, we propose a new framework called the extended feature set for detecting copies of images. In our approach, virtual prior attacks are applied to copyrighted images to generate novel features, which serve as training data. The copy-detection problem can be solved by learning classifiers from the training data, thus, generated. Our approach can be integrated into existing copy detectors to further improve their performance. Experiment results demonstrate that the proposed approach can substantially enhance the accuracy of copy detection.
Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen
IEEE Trans. Image Process.3
2006 Query taxonomy generation for web search
abstract
We propose an approach that organizes the search-result clusters into a hierarchical structure, called a query taxonomy, from the user's perspective. The proposed approach is based on an unsupervised classification method, which uses the dynamic Web as the training corpus. With query taxonomy, users can browse relevant Web documents more conveniently and comprehensibly. Our experimental results verify the feasibility and the effectiveness of the proposed approach to query taxonomy generation in Web search.
Pu-Jen Cheng, Ching-Hsiang Tsai, Chen-Ming Hung, Lee-Feng Chien
CIKM4
2006 Image Copy Detection via Grouping in Feature Space Based on Virtual Prior Attacks
abstract
In the past, many researches on image copy detection focused on finding a feature that is robust enough for various kinds of image attacks. But it is difficult to find a globally effective feature that is appropriate for many situations. In this paper, we introduce a classification framework to this problem, instead of solving the feature-selection problem. In our approach, novel features are generated by applying virtual prior attacks to copyrighted images, and the copy-detection problem is converted to a classification one that is more robust to solve. Our approach can combine existing image copy detectors and further raise their performances.
Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen
ICIP3
2006 Image Content Clustering and Summarization for Photo Collections
abstract
Rapid growth of digital photography in recent years spurred the need of photo management tools. In this study, we propose an automatic organization framework for photo collections based on image content, so that a novel browsing experience is provided for users. For each photograph, human faces, together with corresponding clothes and nearby regions are located. We extract color histograms of these regions as the image content feature. Then a similarity matrix of a photo collection is generated according to temporal and content features of those photographs. We perform hierarchical clustering based on this matrix, and extract duplicate subjects of a cluster by introducing the contrast context histogram (CCH) technique. The experimental results show that the developed framework provides a promising result for photo management
Cheng-Hung Li, Chih-Yi Chiu, Chun-Rong Huang, Chu-Song Chen, Lee-Feng Chien
ICME5
2006 Exploiting the Web as the multilingual corpus for unknown query translation
abstract
Abstract Users' cross‐lingual queries to a digital library system might be short and the query terms may not be included in a common translation dictionary (unknown terms). In this article, the authors investigate the feasibility of exploiting the Web as the multilingual corpus source to translate unknown query terms for cross‐language information retrieval in digital libraries. They propose a Web‐based term translation approach to determine effective translations for unknown query terms by mining bilingual search‐result pages obtained from a real Web search engine. This approach can enhance the construction of a domain‐specific bilingual lexicon and bring multilingual support to a digital library that only has monolingual document collections. Very promising results have been obtained in generating effective translation equivalents for many unknown terms, including proper nouns, technical terms, and Web query terms, and in assisting bilingual lexicon construction for a real digital library system.
Jenq-Haur Wang, Jei-Wen Teng, Wen-Hsiang Lu, Lee-Feng Chien
J. Assoc. Inf. Sci. Technol.4
2005 Annotating Text Segments in Documents for Search
abstract
It has been shown that annotating prominent text patterns contained in documents with appropriate types may benefit many applications. Most conventional tools for automatic text annotation extract named entities from texts and annotate them with information about persons, locations, dates and so on. However, this kind of entity type information is often short in length and is mostly limited to a small set of broader categories. In this paper, we try to remedy this problem by presenting an approach to extract global evidences from documents for improved named entity recognition. We also propose an unsupervised, generalized classification approach that collects training data from the Web automatically and classifies text patterns into more refined categories. Experimental results show the feasibility of the proposed approaches for search on the data of the NTCIR-2 information retrieval task.
Pu-Jen Cheng, Hsin-Chen Chiao, Yi-Cheng Pan, Lee-Feng Chien
Web Intelligence4
2005 Automatic Training Corpora Acquisition through Web Mining
abstract
Text classification is a task having been extensively studied for decades. However, most previous work pre-assumes the existence of explicitly labeled corpora. In this study, we focus on the issue of automatic corpora acquisition. We propose a Web-based mining approach to collect necessary corpora, which can be greatly useful to both common users and system designers. Moreover, the proposed technique can also be incorporated with existing classification techniques to further boost classifier performance. It has been shown that the concept of the class can be captured by the class name and its associated terms (Huang et al., 2004). In this work, we aim at analyzing Web-retrieved documents to discover the associated terms, which are further utilized to collect more training corpora. Working iteratively, the proposed approach can acquire training corpora of high quality. We give empirical evidence that the classifiers thus created have promising accuracy. In sum, the convenience and efficiency of the proposed approach, along with the new perspective on the issue of corpora acquisition, are the primary contributions of this work.
Kuan-Ming Lin, Lee-Feng Chien
Web Intelligence3
2005 Taxonomy generation for text segments: A practical web-based approach
abstract
It is crucial in many information systems to organize short text segments, such as keywords in documents and queries from users, into a well-formed taxonomy. In this article, we address the problem of taxonomy generation for diverse text segments with a general and practical approach that uses the Web as an additional knowledge source. Unlike long documents, short text segments typically do not contain enough information to extract reliable features. This work investigates the possibilities of using highly ranked search-result snippets to enrich the representation of text segments. A hierarchical clustering algorithm is then designed for creating the hierarchical topic structure of text segments. Text segments with close concepts can be grouped together in a cluster, and relevant clusters linked at the same or near levels. Different from traditional clustering algorithms, which tend to produce cluster hierarchies with a very unnatural shape, the algorithm tries to produce a more natural and comprehensive tree hierarchy. Extensive experiments were conducted on different domains of text segments, including subject terms, people names, paper titles, and natural language questions. The obtained experimental results have shown the potential of the proposed approach, which provides a basis for the in-depth analysis of text segments on a larger scale and is believed able to benefit many information systems.
Shui-Lung Chuang, Lee-Feng Chien
ACM Trans. Inf. Syst.2
2004 Creating Multilingual Translation Lexicons with Regional Variations Using Web Corpora
abstract
The purpose of this paper is to automatically create multilingual translation lexicons with regional variations. We propose a transitive translation approach to determine translation variations across languages that have insufficient corpora for translation via the mining of bilingual search-result pages and clues of geographic information obtained from Web search engines. The experimental results have shown the feasibility of the proposed approach in efficiently generating translation equivalents of various terms not covered by general translation dictionaries. It also revealed that the created translation lexicons can reflect different cultural aspects across regions such as Taiwan, Hong Kong and mainland China.
Pu-Jen Cheng, Wen-Hsiang Lu, Jei-Wen Teng, Lee-Feng Chien
ACL4
2004 A practical web-based approach to generating topic hierarchy for text segments
abstract
It is crucial in many information systems to organize short text segments, such as keywords in documents and queries from users, into a well-formed topic hierarchy. In this paper, we address the problem of generating topic hierarchies for diverse text segments with a general and practical approach that uses the Web as an additional knowledge source. Unlike long documents, short text segments typically do not contain enough information to extract reliable features. This work investigates the possibilities of using highly ranked search-result snippets to enrich the representation of text segments. A hierarchical clustering algorithm is then applied to create the hierarchical topic structure of text segments. Different from traditional clustering algorithms, which tend to produce cluster hierarchies with a very unnatural shape, the approach tries to produce a more natural and comprehensive hierarchy. Extensive experiments were conducted on different domains of text segments. The obtained results have shown the potential of the proposed approach, which is believed able to benefit many information systems.
Shui-Lung Chuang, Lee-Feng Chien
CIKM2
2004 Mining the Web for Generating Thematic Metadata from Textual Data
abstract
Conventional tools for automatic metadata creation mostly extract named entities or patterns from texts and annotate them with information about persons, locations, dates, and so on. However, this kind of entity type information is often too primitive for more advanced intelligent applications such as concept-based search. Here, we try to generate semantically-deep metadata with limited human intervention. The main idea behind our approach is to use Web mining and categorization techniques to create thematic metadata. The proposed approach, comprises of three computational modules: feature extraction, HCQF (hier-concept query formulation) and text instance categorization. The feature extraction module sends the name of text instances to Web search engines, and the returned highly-ranked search-result pages are used to describe them.
Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien
ICDE3
2004 Categorizing Unknown Text Segments for Information Extraction Using a Search Result Mining Approach
Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien
IJCNLP3
2004 Generating Concept Hierarchies from Text for Intelligence Analysis
Jenq-Haur Wang, Chien-Chung Huang 0004, Jei-Wen Teng, Lee-Feng Chien
ISI4
2004 Translating unknown queries with web corpora for cross-language information retrieval
abstract
It is crucial for cross-language information retrieval (CLIR) systems to deal with the translation of unknown queries due to that real queries might be short. The purpose of this paper is to investigate the feasibility of exploiting the Web as the corpus source to translate unknown queries for CLIR. We propose an online translation approach to determine effective translations for unknown query terms via mining of bilingual search-result pages obtained from Web search engines. This approach can alleviate the problem of the lack of large bilingual corpora, translate many unknown query terms, provide flexible query specifications, and extract semantically-close translations to benefit CLIR tasks -- especially for cross-language Web search.
Pu-Jen Cheng, Jei-Wen Teng, Ruey-Cheng Chen, Jenq-Haur Wang, Wen-Hsiang Lu, Lee-Feng Chien
SIGIR6
2004 Liveclassifier: creating hierarchical text classifiers through web corpora
abstract
Many Web information services utilize techniques of information extraction(IE) to collect important facts from the Web. To create more advanced services, one possible method is to discover thematic information from the collected facts through text classification. However, most conventional text classification techniques rely on manual-labelled corpora and are thus ill-suited to cooperate with Web information services with open domains. In this work, we present a system named LiveClassifier that can automatically train classifiersthrough Web corpora based on user-defined topic hierarchies. Due to its flexibility and convenience, LiveClassifier can be easily adapted for various purposes. New Web information services can be created to fully exploit it; human users can use it to create classifiers for their personal applications. The effectiveness of classifiers created by LiveClassifier is well supportedby empirical evidence.
Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien
WWW3
2004 Using a web-based categorization approach to generate thematic metadata from texts
abstract
Conventional tools for automatic metadata creation mostly extract named entities or text segments from texts and annotate them with information about persons, locations, dates, and so on. However, this kind of entity type information is often insufficient for machines to understand the facts contained in the texts, thus precluding the possibility of implementing more advanced, intelligent applications, such as concept-based search. In this work, we try to create more refined thematic metadata inherent in texts. Based on Web resource mining, our approach acquires training corpora necessary to describe both the thematic categories and the metadata extracted from the texts. The approach then finds the corresponding relationships among them by means of categorization and thus generates thematic metadata for the textual data. Experimental results confirm the potential and wide adaptability of our approach.
Chien-Chung Huang 0004, Shui-Lung Chuang, Lee-Feng Chien
ACM Trans. Asian Lang. Inf. Process.3
2004 Anchor text mining for translation of Web queries: A transitive translation approach
abstract
To discover translation knowledge in diverse data resources on the Web, this article proposes an effective approach to finding translation equivalents of query terms and constructing multilingual lexicons through the mining of Web anchor texts and link structures. Although Web anchor texts are wide-scoped hypertext resources, not every particular pair of languages contains sufficient anchor texts for effective extraction of translations for Web queries. For more generalized applications, the approach is designed based on a transitive translation model. The translation equivalents of a query term can be extracted via its translation in an intermediate language. To reduce interference from translation errors, the approach further integrates a competitive linking algorithm into the process of determining the most probable translation. A series of experiments has been conducted, including performance tests on term translation extraction, cross-language information retrieval, and translation suggestions for practical Web search services, respectively. The obtained experimental results have shown that the proposed approach is effective in extracting translations of unknown queries, is easy to combine with the probabilistic retrieval model to improve the cross-language retrieval performance, and is very useful when the considered language pairs lack a sufficient number of anchor texts. Based on the approach, an experimental system called LiveTrans has been developed for English--Chinese cross-language Web search.
Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee
ACM Trans. Inf. Syst.2
2003 Auto-generation of topic hierarchies for web images from users' perspectives
abstract
In this paper, we propose an approach to automatically generating a Yahoo!-like topic hierarchy for organizing Web images from users' perspectives. Relatively little effort has been devoted towards providing such a taxonomy simultaneously considering users' image requests for semantic and visual information. Based on the characteristic that a Web-image query may be refined by various attributes, the proposed approach hierarchically groups similar queries from search engine logs into topic classes at different semantic levels. The generated topic hierarchy has the advantages of organizing image data from users' perspectives for browsing, searching, annotation and users' needs analysis.A series of experiments have been conducted on real-world image search engine logs. Experimental results show that the proposed approach is feasible to generate topic hierarchies for Web images. Moreover, the generated hierarchy has been successfully applied to analysis of users' search interests, which have more focuses on some specific domains when compared with document requests.
Pu-Jen Cheng, Lee-Feng Chien
CIKM2
2003 Enriching Web taxonomies through subject categorization of query terms from search engine logs
Shui-Lung Chuang, Lee-Feng Chien
Decis. Support Syst.2
2003 Relevant term suggestion in interactive web search based on contextual information in query session logs
abstract
Abstract This paper proposes an effective term suggestion approach to interactive Web search. Conventional approaches to making term suggestions involve extracting co‐occurring keyterms from highly ranked retrieved documents. Such approaches must deal with term extraction difficulties and interference from irrelevant documents, and, more importantly, have difficulty extracting terms that are conceptually related but do not frequently co‐occur in documents. In this paper, we present a new, effective log‐based approach to relevant term extraction and term suggestion. Using this approach, the relevant terms suggested for a user query are those that co‐occur in similar query sessions from search engine logs, rather than in the retrieved documents. In addition, the suggested terms in each interactive search step can be organized according to its relevance to the entire query session, rather than to the most recent single query as in conventional approaches. The proposed approach was tested using a proxy server log containing about two million query transactions submitted to search engines in Taiwan. The obtained experimental results show that the proposed approach can provide organized and highly relevant terms, and can exploit the contextual information in a user's query session to make more effective suggestions.
Chien-Kang Huang, Lee-Feng Chien, Yen-Jen Oyang
J. Assoc. Inf. Sci. Technol.2
2002 A Transitive Model for Extracting Translation Equivalents of Web Queries through Anchor Text Mining
Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee
COLING2
2002 Towards Automatic Generation of Query Taxonomy: A Hierarchical Query Clustering Approach
abstract
Most previous work on automatic query clustering generated a flat, un-nested partition of query terms. In this work, we discuss the organization of query terms into a hierarchical structure and construct a query taxonomy in an automatic way. The proposed approach is designed based on a hierarchical agglomerative clustering algorithm to hierarchically group similar queries and generate cluster hierarchies using a novel cluster partition technique. The search processes of real-world search engines are combined to obtain highly ranked Web documents as the feature source for each query term. Preliminary experiments show that the proposed approach is effective for obtaining thesaurus information for query terms, and is also feasible for constructing a query taxonomy which provides a basis for in-depth analysis of users' search interests and domain-specific vocabulary on a larger scale.
Shui-Lung Chuang, Lee-Feng Chien
ICDM2
2002 Incremental Extraction of Keyterms for Classifying Multilingual Documents in the Web
Lee-Feng Chien, Chien-Kang Huang, Hsin-Chen Chiao, Shih-Jui Lin
PAKDD1
2002 Translation of web queries using anchor text mining
abstract
This article presents an approach to automatically extracting translations of Web query terms through mining of Web anchor texts and link structures. One of the existing difficulties in cross-language information retrieval (CLIR) and Web search is the lack of appropriate translations of new terminology and proper names. The proposed approach successfully exploits the anchor-text resources and reduces the existing difficulties of query term translation. Many query terms that cannot be obtained in general-purpose translation dictionaries are, therefore, extracted.
Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee
ACM Trans. Asian Lang. Inf. Process.2
2001 Anchor Text Mining for Translation of Web Queries
abstract
The paper presents an approach to automatically extracting translations of Web query terms through mining of Web anchor texts and link structures. One of the existing difficulties in cross-language information retrieval (CLIR) and Web search is the lack of the appropriate translations of new terminology and proper names. Such a difficult problem can be effectively alleviated by our proposed approach, and the resource of anchor texts in the Web is proven a valuable corpus for this kind of term translation.
Wen-Hsiang Lu, Lee-Feng Chien, Hsi-Jian Lee
ICDM2
2001 Anchor Text Mining for Translation Extraction of Query Terms
abstract
This paper presents an approach to automatically extracting the bilingual translations of many Web query terms through mining the Web anchor texts. Some preliminary experiments are conducted on using 109,416 Web pages containing both Chinese and English anchor texts in their in-links to extract Chinese translations of 200 English queries selected from popular query terms in Taiwan. It is found that the effective translations of 75% of the popular query terms can be extracted, in which 87.2% cannot be obtained in common translation dictionaries.
Wen-Hsiang Lu, Hsi-Jian Lee, Lee-Feng Chien
SIGIR3
2001 A Contextual Term Suggestion Mechanism for Interactive Web Search
Chien-Kang Huang, Yen-Jen Oyang, Lee-Feng Chien
Web Intelligence3
2000 Live thesaurus construction for interactive voice-based web search
abstract
Since Web users ’ queries are often too short, an accurate and interactive speech interface is believed very helpful, especially for WAP-based Web search engines. To provide high accurate speech recognition and effective interactive search, a rigid and live web thesaurus that contains users ' search terms plus a set of relations between their associated terms is highly in demand. The purpose of this paper is intended to present a log-based approach for live thesaurus construction. Based on the live thesaurus, certain kinds of users ' information behaviors could be characterized and a more effective voice-based search engine could be developed.
Shui-Lung Chuang, Hsiao-Tieh Pu, Wen-Hsiang Lu, Lee-Feng Chien
INTERSPEECH4
2000 Live Lexicons and Dynamic Corpora Adapted to the Network Resources for Chinese Spoken Language Processing Applications in an Internet Era
Lin-Shan Lee, Lee-Feng Chien
LREC2
2000 Auto-construction of a live thesaurus from search term logs for interactive Web search
abstract
The purpose of this paper is to present an on-going research that is intended to construct a live thesaurus directly from search term logs of real-world search engines. Such a thesaurus designed can contain representative search terms, their frequency in use, the corresponding subject categories, the associated and relevant terms, and the hot visiting Web sites/pages the search terms may reach.
Shui-Lung Chuang, Hsiao-Tieh Pu, Wen-Hsiang Lu, Lee-Feng Chien
SIGIR4
2000 A spoken-access approach for chinese text and speech information retrieval
abstract
This paper presents an efficient spoken-access approach for both Chinese text and Mandarin speech information retrieval. The proposed approach is developed not only to deal with the retrieval of spoken documents, but also to improve the capability of human-computer interaction via voice input for information-retrieval systems. Based on utilization of the monosyllabic structure of the Chinese language, the proposed approach can tolerate speech recognition errors by performing speech query recognition and approximate information retrieval at the syllable-level. Furthermore, with the help of automatic term suggestion and relevance feedback techniques, the proposed approach is robust in enabling users using voice input to interact with IR systems at each stage of the retrieval process. Extensive experiments show that the proposed approach can improve the effectiveness of information retrieval via speech interaction. The encouraging results suggest that a Mandarin speech interface for information retrieval and digital library systems can, therefore, be developed.
Lee-Feng Chien, Hsin-Min Wang, Bo-Ren Bai, Sun-Chien Lin
J. Am. Soc. Inf. Sci.1
1999 An OODBMS-IRS Integration Based on a Statistical Corpus Extraction Method for Document Management
Chung-Hong Lee, Lee-Feng Chien
DEXA2
1999 PAT-tree-based adaptive keyphrase extraction for intelligent Chinese information retrieval
Lee-Feng Chien
Inf. Process. Manag.1
1998 A Novel Integration of OODBMS and Information Retrieval Techniques for a Document Repository
Chung-Hong Lee, Lee-Feng Chien
DEXA2
1998 Statistics-based segment pattern lexicon-a new direction for Chinese language modeling
abstract
This paper presents a new direction for Chinese language modeling based on a different concept of the lexicon. Because every Chinese character has its own meaning and there are no "blanks" in Chinese sentences serving as word boundaries, also because the wording structure in the Chinese language is extremely flexible, the "words" in Chinese are actually not well defined, and there does not exist a commonly accepted lexicon. This makes language modeling very sophisticated in the Chinese language, and the "out of vocabulary (OOV)" problem specially serious. A new concept for the lexicon is thus proposed. The elements of this lexicon can be words or any other "segment patterns". They should be extracted from the training corpus by statistical approaches with a goal to minimize the overall perplexity. The language models can then be developed based on this new lexicon. Very encouraging experimental results have been obtained.
Kae-Cherng Yang, Tai-Hsuan Ho, Lee-Feng Chien, Lin-Shan Lee
ICASSP3
1998 A*-admissible key-phrase spotting with sub-syllable level utterance verification
abstract
In this paper, we propose an A*-admissible key-phrase spotting framework, which needs little domain knowledge and is capable of extracting salient key-phrase fragments from an input utterance in real-time. There are two key features in our approach. Firstly, the acoustic models and the search framework are specially designed such that very high degree vocabulary flexibility can be achieved for any desired application tasks. Secondly, the search framework uses an efficient two-pass A* search to generate N-best key-phrase candidates and then several sub-syllable level verification functions are properly weighted and used to further improve the recognition accuracy. Experimental results show that the A*-admissible key-phrase spotting with sub-word level utterance method outperforms the baseline methods used in common approaches. 1. INTRODUCTION In recent years, various spoken dialog systems have been widely investigated for the fast growing demand for real-world applications. It is diff...
Berlin Chen, Hsin-Min Wang, Lee-Feng Chien, Lin-Shan Lee
ICSLP3
1998 Automatic Acquisition of Phrasal Knowledge for English-Chinese Bilingual Information Retrieval
abstract
No abstract available.
Ming-Jer Lee, Lee-Feng Chien
SIGIR2
1997 Internet Chinese information retrieval using unconstrained Mandarin speech queries based on a client-server architecture and a PAT-tree-based language model
abstract
In order to pursue high performance of Chinese information access on the Internet, this paper presents an attractive approach with a successful integration of efficient speech recognition and information retrieval techniques. A working system based on the proposed approach for speech retrieval of real-time Chinese netnews services has been implemented and tested. Very exciting performance has been achieved.
Lee-Feng Chien, Sung-Chien Lin, Jenn-Chau Hong, Ming-Chiuan Chen, Hsin-Min Wang, Jia-Lin Shen, Keh-Jiann Chen, Lin-Shan Lee
ICASSP1
1997 Syllable-based relevance feedback techniques for Mandarin voice record retrieval using speech queries
abstract
In order to solve the problem with the new environment of fast growth of audio resources on the Internet, we have presented a syllable-based approach which is capable of retrieving Mandarin voice records using queries of unconstrained speech. However, the performance achieved by this previously proposed approach is still not satisfactory, and one of the reason is that very often the information provided by the speech query for the request subject may not be sufficient. We present approaches based the relevance feedback technique to improving the performance of the previous research. The proposed approaches include a relevance measure adjustment scheme using a relevance table for the voice database, a query expansion scheme to generate a new query including the feedback information, and a combination of these two schemes. Extensive preliminary experiments were performed and demonstrated.
Lin-Shan Lee, Bo-Ren Bai, Lee-Feng Chien
ICASSP3
1997 Intelligent retrieval of very large Chinese dictionaries with speech queries
abstract
To retrieve a Chinese word from a Chinese dictionary, it needs the user to know exactly the first character of the desired word. Because there is more than 10,000 Chinese characters, this makes the Chinese dictionary relatively difficult to be used. To reduce the problem, this paper presents intelligent retrieval techniques for very large Chinese dictionaries with speech queries. The proposed techniques properly integrate the technologies of Mandarin speech recognition and Chinese information retrieval with a syllable-based approach utilizing the mono-syllabic structure of the language. Moreover, it is very nice to provide the function of retrieving all relevant word entries from the dictionaries using speech queries describing “general concepts” of the desired words. To achieve the challenging function, the techniques of relevance feedback are also included. Based on these techniques, a retrieval system was implemented successfully on a Pentium PC for a very large Chinese dictionary which includes 160,000 word entries and the total length of the lexical information under the word entries exceeds 20,000,000 words.
Sung-Chien Lin, Lee-Feng Chien, Ming-Chiuan Chen, Lin-Shan Lee, Keh-Jiann Chen
EUROSPEECH2
1997 Chinese language model adaptation based on document classification and multiple domain-specific language models
Sung-Chien Lin, Chi-Lung Tsai, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH3
1997 PAT-tree-based Keyword Extraction for Chinese Information Retrieval
abstract
urgent need to promote Chinese in this paper we will raise the significance of keyword extraction using a new PAT-treebased approach, which is efficient in automatic keyword extraction from a set of relevant Chinese documents.This approach has been successfully applied in several IR researches, such as document classification, book indexing and relevance feedback.Many Chinese language processing applications therefore step ahead from character level to word/phrase level,
Lee-Feng Chien
SIGIR1
1996 An efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a clustered language model
abstract
This paper presents an accurate and efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a specially-designed clustered language model. To reduce the problems resulted from the complexity of unconstrained speech-input queries for retrieval, the system is completely syllable-based in both speech recognition and database retrieval by properly utilizing the mono-syllabic structure of Chinese language. In addition, it partitions the records in the database into clusters and trains the clustered language model using the clustering results. The proposed clustered language model with its augmented search algorithm are very useful to improve accuracy and speed of the speech retrieval system. In the preliminary tests using an experimental database with about 30,000 bibliographical records, it was found that the present system can accept unconstrained speech-input queries and achieve very good performance.
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICASSP2
1996 Very-large-vocabulary Mandarin voice message file retrieval using speech queries
abstract
In order to solve the problem with the new environment of fast growth of audio resources on the Intemet, this paper presents a new approach which is capable of retrieving Mandarin voice message files using queries of unconstrained speech.By properly utilizing the monosyllabic structure of the Chinese language, the proposed approach perfoms the statistical similarity estimation between the speech queries and the voice message files, and executes the complete matching process directly at the phonetic level using syllable-based statistical information.Based on this approach, some experiments are tested and encouraging results are demonstrated.
Bo-Ren Bai, Lee-Feng Chien, Lin-Shan Lee
ICSLP2
1996 Speaker intention modeling for large vocabulary Mandarin spoken dialogues
abstract
This paper presents a statistical speaker intention modeling approach of speech act types (SAT's)[l] prediction for large vocabulary Mandarin spoken dialogues.A SAT is an abstraction of speaker's intention in terms of the type of action thax the speaker intends by the utterance.With this approach, spoken dialogue systems can be constructed to predict speaker's intention and make a proper action in advance.
Yen-Ju Yang, Lee-Feng Chien, Lin-Shan Lee
ICSLP2
1995 Golden Mandarin (III)-a user-adaptive prosodic-segment-based Mandarin dictation machine for Chinese language with very large vocabulary
abstract
This paper presents a prototype prosodic-segment-based Mandarin dictation machine for the Chinese language with very large vocabulary. It accepts utterances continuous within a prosodic segment which is composed of one or a few word(s). It also possesses various on-line learning capabilities for fast adaptation to a new user in acoustic, lexical and linguistic levels. The overall system is implemented on an IBM/PC with an additional DSP card including a Motorola DSP 96002 chip. The word accuracy can achieve nearly 90% for a new user after he produces about 10 minutes of speech to train the system, and the accuracy can be further improved with the on-line learning functions.
Ren-Yuan Lyu, Lee-Feng Chien, Shiao-Hong Hwang, Hung-Yun Hsieh, Rung-Chiuan Yang, Bo-Ren Bai, Jia-Chi Weng, Yen-Ju Yang, Shi-Wei Lin, Keh-Jiann Chen, Chiu-yu Tseng, Lin-Shan Lee
ICASSP2
1995 Fast and accurate continuous speech recognition for Chinese language with very large vocabulary
Tai-Hsuan Ho, Hsin-Min Wang, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH3
1995 A syllable-based very-large-vocabulary voice retrieval system for Chinese databases with textual attributes
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH2
1995 Unconstrained speech retrieval for Chinese document databases with very large vocabulary and unlimited domains
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH2
1995 Fast and Quasi-Natural Language Search for Gigabits of Chinese Texts
abstract
Article Fast and quasi-natural language search for gigabytes of Chinese texts Share on Author: Lee-Feng Chien Institute of Information Science, Academia Sinica, Taipei, Taiwan, R.O.C. Institute of Information Science, Academia Sinica, Taipei, Taiwan, R.O.C.View Profile Authors Info & Claims SIGIR '95: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrievalJuly 1995 Pages 112–120https://doi.org/10.1145/215206.215345Online:01 July 1995Publication History 26citation508DownloadsMetricsTotal Citations26Total Downloads508Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Lee-Feng Chien
SIGIR1
1994 An intelligent and efficient word-class-based Chinese language model for Mandarin speech recognition with very large vocabulary
Yen-Ju Yang, Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICSLP3
1993 Golden Mandarin (II)-an improved single-chip real-time Mandarin dictation machine for Chinese language with very large vocabulary
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen, I-Jung Hung, Ming-Yu Lee, Lee-Feng Chien, Yumin Lee, Ren-Yuan Lyu, Hsin-Min Wang, Yung-Chuan Wu, Tung-Sheng Lin, Hung-Yan Gu, Chi-ping Nee, Chun-Yi Liao, Yeng-Ju Yang, Yuan-Cheng Chang, Rung-Chiung Yang
ICASSP (2)6
1993 A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications
abstract
A language processing model is proposed in which the grammatical approach of unification grammar and the statistical approach of Markov language models are properly integrated in a word lattice chart parsing algorithm with different best-first parsing strategies. This model has been successfully implemented in experiments on Mandarin speech recognition although it is language-independent. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved. A correct rate of 93.8% and 5 s per sentence on an IBM PC/AT, as compared with 73.8% and 25 s using unification grammar alone and 82.2% and 3 s using a Markov language model alone, was achieved. This high performance is due to the effective rejection of noisy word hypothesis interferences; that is, the unification-based grammatical analysis eliminates all illegal combinations, while the Markovian probabilities of constituents combined with the considerations on constituent length indicate the correct direction of processing.>
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
IEEE Trans. Speech Audio Process.1
1991 A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications
abstract
A language processor is to find out a most promising sentence hypothesis for a given word lattice obtained from acoustic signal recognition. In this paper a new language processor is proposed, in which unification grammar and Markov language model are integrated in a word lattice parsing algorithm based on an augmented chart, and the island-driven parsing concept is combined with various preference-first parsing strategies defined by different construction principles and decision rules. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved.
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ACL1
1991 An Efficient Natural Language Processing System Specially Designed for the Chinese Language
Lin-Shan Lee, Lee-Feng Chien, Long Ji Lin, Keh-Jiann Chen
Comput. Linguistics2
1991 An augmented chart data structure with efficient word lattice parsing scheme in speech recognition applications
Lee-Feng Chien, Lin-Shan Lee, Keh-Jiann Chen
Speech Commun.1
1990 An Augmented Chart Data Structure with Efficient Word Lattice Parsing Scheme In Speech Recognition Applications
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
COLING1
1990 An augmented chart parsing algorithm integrating unification grammar and Markov language model for continuous speech recognition
abstract
An efficient algorithm is developed to handle the difficulties in parsing noise word lattices (sets of word hypotheses obtained in continuous-speech recognition) which include problems such as word boundary overlapping, homonyms, lexical ambiguities, recognition uncertainty and errors, etc. An augmented chart is proposed, and the algorithms is then derived on this chart. This algorithm properly integrates the global structural synthesis capabilities of the unification grammar and the local relation estimation capabilities of the Markov language model. The parsing algorithm is island driven and best first. In this way, the features of the grammatical and statistical approaches can be combined, and the effects of the two different approaches are reflected in a single algorithm such that the overall selectivity can be appropriately optimized.>
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICASSP1