Jinxi Xu

dblp:67/1760 · DBLP profile ↗
← Back
18ranked-venue papers
12as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 9 first-authorArtificial intelligence and machine learning · 7 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Information retrieval · 95% Data mining · 5%
Artificial intelligence
5 papers
Machine translation · 76% Language models and text generation · 12% Question answering and dialogue systems · 7%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
statistical machine translation
0.332010
Statistical Machine Translation with a Factorized Grammar · EMNLP 2010
Effective Use of Linguistic and Contextual Information for Statistical Machine Translation · EMNLP 2009
A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008
Information retrieval › query reformulation
query expansion
0.352013
Query expansion using path-constrained random walks · SIGIR 2013
Empirical studies in strategies for Arabic retrieval · SIGIR 2002
Improving the effectiveness of information retrieval with local context analysis · ACM Trans. Inf. Syst. 2000
Natural language and speech › Machine translation
syntax-based machine translation
0.122010
A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008
Statistical Machine Translation with a Factorized Grammar · EMNLP 2010
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.112010
Combining Unsupervised and Supervised Alignments for MT: An Empirical Study · EMNLP 2010
Information retrieval
cross-language information retrieval
0.132002
Empirical studies in strategies for Arabic retrieval · SIGIR 2002
Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001
Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000
Natural language and speech › Language models and text generation › language modeling
dependency language model
0.112008
A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008
Information retrieval
retrieval models
0.132001
Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001
Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999
Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000
Information retrieval
query log analysis
0.012013
Query expansion using path-constrained random walks · SIGIR 2013
Information retrieval
web search
0.012013
Query expansion using path-constrained random walks · SIGIR 2013
Natural language and speech › Question answering and dialogue systems › open-ended question answering
definitional question answering
0.012004
Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004
Data mining › text mining
information extraction and text analysis
0.012004
Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004
Information retrieval
distributed information retrieval
0.021999
Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999
Effective Retrieval with Distributed Collections · SIGIR 1998
Information retrieval › distributed information retrieval
resource selection
0.021999
Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999
Effective Retrieval with Distributed Collections · SIGIR 1998
Information retrieval › retrieval models
probabilistic retrieval model
0.022001
Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001
Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000
Information retrieval › cross-language information retrieval
arabic retrieval
0.012002
Empirical studies in strategies for Arabic retrieval · SIGIR 2002
Information retrieval › document retrieval › concept-based retrieval
thesaurus-based retrieval
0.012002
Empirical studies in strategies for Arabic retrieval · SIGIR 2002
Information retrieval › query reformulation › query expansion
local context analysis
0.012000
Improving the effectiveness of information retrieval with local context analysis · ACM Trans. Inf. Syst. 2000
Information retrieval › cross-language information retrieval
query translation
0.012000
Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000
Information retrieval › retrieval models › language model
cluster-based language model
0.011999
Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999
Information retrieval › retrieval models
language model
0.011999
Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999
Information retrieval › indexing
stemming
0.011998
Corpus-Based Stemming Using Cooccurrence of Word Variants · ACM Trans. Inf. Syst. 1998
Information retrieval
evaluation
0.012004
Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004

Methods — techniques the papers use, named apart from their topics

random walk · 0.2learning to rank · 0.2factorized grammar · 0.1empirical study · 0.1statistical machine translation · 0.1information extraction · 0.1ROUGE · 0.1target dependency language model · 0.1dependency parsing · 0.1stemming · 0.0spelling normalization · 0.0probabilistic term translation · 0.0character tri-grams · 0.0machine translation · 0.0generative model · 0.0
YearPublicationVenuePosition
2013 Query expansion using path-constrained random walks
abstract
This paper exploits Web search logs for query expansion (QE) by presenting a new QE method based on path-constrained random walks (PCRW), where the search logs are represented as a labeled, directed graph, and the probability of picking an expansion term for an input query is computed by a learned combination of constrained random walks on the graph. The method is shown to be generic in that it covers most of the popular QE models as special cases and flexible in that it provides a principled mathematical framework in which a wide variety of information useful for QE can be incorporated in a unified way. Evaluation is performed on the Web document ranking task using a real-world data set. Results show that the PCRW-based method is very effective for the expansion of rare queries, i.e., low-frequency queries that are unseen in search logs, and that it outperforms significantly other state-of-the-art QE meth-ods.
Jianfeng Gao 0001, Gu Xu, Jinxi Xu
SIGIR3
2010 Statistical Machine Translation with a Factorized Grammar
Libin Shen, Bing Zhang 0004, Spyridon Matsoukas, Jinxi Xu, Ralph M. Weischedel
EMNLP4
2010 Combining Unsupervised and Supervised Alignments for MT: An Empirical Study
Jinxi Xu, Antti-Veikko I. Rosti
EMNLP1
2010 String-to-Dependency Statistical Machine Translation
abstract
We propose a novel string-to-dependency algorithm for statistical machine translation. This algorithm employs a target dependency language model during decoding to exploit long distance word relations, which cannot be modeled with a traditional n-gram language model. Experiments show that the algorithm achieves significant improvement in MT performance over a state-of-the-art hierarchical string-to-string system on NIST MT06 and MT08 newswire evaluation sets.
Libin Shen, Jinxi Xu, Ralph M. Weischedel
Comput. Linguistics2
2009 Effective Use of Linguistic and Contextual Information for Statistical Machine Translation
Libin Shen, Jinxi Xu, Bing Zhang 0004, Spyridon Matsoukas, Ralph M. Weischedel
EMNLP2
2008 A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model
Libin Shen, Jinxi Xu, Ralph M. Weischedel
ACL2
2005 Empirical studies on the impact of lexical resources on CLIR performance
Jinxi Xu, Ralph M. Weischedel
Inf. Process. Manag.1
2004 Evaluation of an extraction-based approach to answering definitional questions
abstract
This paper evaluates an extraction-based approach to answering definitional questions. Our system extracted useful linguistic constructs called linguistic features from raw text using information extraction tools and formulated answers based on such features. The features employed include appositives, copulas, structured patterns, relations, propositions and raw sentences. The features were ranked based on feature type and similarity to a question profile. Redundant features were detected using a simple heuristic-based strategy. The approach achieved state of the art performance at the TREC 2003 QA evaluation. Component analysis of the system was carried out using an automatic scoring function called Rouge (Lin and Hovy, 2003). Major findings include 1) answers using linguistic features are significantly better than those using raw sentences; 2) the most useful features are appositives and copulas; 3) question profiles, as a means of modeling user interests, can significantly improve system performance; 4) the Rouge scores are closely correlated with subjective evaluation results, indicating the suitability of using Rouge for evaluating definitional QA systems.
Jinxi Xu, Ralph M. Weischedel, Ana Licuanan
SIGIR1
2003 Cross-lingual retrieval for Hindi
abstract
In this paper we describe the evaluation results of applying a cross-lingual retrieval model to retrieve Hindi documents relevant to an English query. Though the technique has been previously applied and evaluated for retrieving Chinese and Arabic documents given an English query, what is new about these experiments is porting the model to Hindi in two weeks' time.
Jinxi Xu, Ralph M. Weischedel
ACM Trans. Asian Lang. Inf. Process.1
2002 Audio Indexing of Arabic broadcast news
abstract
This paper describes the development of the BBN Audio Indexing System for broadcast news in Arabic. Key issues addressed in this work revolve around the three major components of the audio indexing system: automatic speech recognition, speaker identification, and named entity identification. The system deals with several challenges introduced by the Arabic language, including the absence of short vowels in written text and the presence of compound words that are formed by the concatenation of certain conjunctions, prepositions, articles, and pronouns, as prefixes and suffixes to the word stem. The lack of short vowels in the transcripts prompted a novel solution that further demonstrated the power of hidden Markov models to deal with ambiguity. Another challenge was the acquisition of appropriate language modeling data, given the absence of broadcast news data for that purpose. We present performance results for all three components of the Audio Indexing System, which we believe represent the state of the art for Arabic broadcast news.
Jayadev Billa, Mohammed Noamany, Amit Srivastava, Daben Liu, Rebecca Stone, Jinxi Xu, John Makhoul, Francis Kubala
ICASSP6
2002 Empirical studies in strategies for Arabic retrieval
abstract
This work evaluates a few search strategies for Arabic monolingual and cross-lingual retrieval, using the TREC Arabic corpus as the test-bed. The release by NIST in 2001 of an Arabic corpus of nearly 400k documents with both monolingual and cross-lingual queries and relevance judgments has been a new enabler for empirical studies. Experimental results show that spelling normalization and stemming can significantly improve Arabic monolingual retrieval. Character tri-grams from stems improved retrieval modestly on the test corpus, but the improvement is not statistically significant. To further improve retrieval, we propose a novel thesaurus-based technique. Different from existing approaches to thesaurus-based retrieval, ours formulates word synonyms as probabilistic term translations that can be automatically derived from a parallel corpus. Retrieval results show that the thesaurus can significantly improve Arabic monolingual retrieval. For cross-lingual retrieval (CLIR), we found that spelling normalization and stemming have little impact.
Jinxi Xu, Alexander Fraser 0001, Ralph M. Weischedel
SIGIR1
2001 Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval
abstract
This work proposes and evaluates a probabilistic cross-lingual retrieval system. The system uses a generative model to estimate the probability that a document in one language is relevant, given a query in another language. An important component of the model is translation probabilities from terms in documents to terms in a query. Our approach is evaluated when 1) the only resource is a manually generated bilingual word list, 2) the only resource is a parallel corpus, and 3) both resources are combined in a mixture model. The combined resources produce about 90% of monolingual performance in retrieving Chinese documents. For Spanish the system achieves 85% of monolingual performance using only a pseudo-parallel Spanish-English corpus. Retrieval results are comparable with those of the structural query translation technique (Pirkola, 1998) when bilingual lexicons are used for query translation. When parallel texts in addition to conventional lexicons are used, it achieves better retrieval results but requires more computation than the structural query translation technique. It also produces slightly better results than using a machine translation system for CLIR, but the improvement over the MT system is not significant.
Jinxi Xu, Ralph M. Weischedel
SIGIR1
2000 Cross-lingual Information Retrieval Using Hidden Markov Models
abstract
This paper presents empirical results in cross-lingual information retrieval using English queries to access Chinese documents (TREC-5 and TREC-6) and Spanish documents (TREC-4). Since our interest is in languages where resources may be minimal, we use an integrated probabilistic model that requires only a bilingual dictionary as a resource. We explore how a combined probability model of term translation and retrieval can reduce the effect of translation ambiguity. In addition, we estimate an upper bound on performance, if translation ambiguity were a solved problem. We also measure performance as a function of bilingual dictionary size.
Jinxi Xu, Ralph M. Weischedel
EMNLP1
2000 Improving the effectiveness of information retrieval with local context analysis
abstract
Techniques for automatic query expansion have been extensively studied in information research as a means of addressing the word mismatch between queries and documents. These techniques can be categorized as either global or local. While global techniques rely on analysis of a whole collection to discover word relationships, local techniques emphasize analysis of the top-ranked documents retrieved for a query. While local techniques have shown to be more effective that global techniques in general, existing local techniques are not robust and can seriously hurt retrieved when few of the retrieval documents are relevant. We propose a new technique, called local context analysis, which selects expansion terms based on cooccurrence with the query terms within the top-ranked documents. Experiments on a number of collections, both English and non-English, show that local context analysis offers more effective and consistent retrieval results.
Jinxi Xu, W. Bruce Croft
ACM Trans. Inf. Syst.1
1999 Cluster-Based Language Models for Distributed Retrieval
abstract
Effective retrieval in a distributed environment is an important but difficult problem. Lack of effectiveness appears to have three causes. First, collection selection based on word histograms is not appropriate for heterogeneous collections. Second, relevant documents are scattered over many collections and searching a few collections misses many relevant documents. Third, most existing collection selection metrics lack sound theoretical justifications and hence may not be well tuned to the problem. We propose a new approach to distributed retrieval based on document clustering and language modeling. Document clustering is used to organize collections around topics. Language modeling is used to properly represent topics and effectively select the right topics for a query. Based on these ideas, three methods are proposed to suit different environments. We show that all three methods improve effectiveness of distributed retrieval. 1 Introduction Information has become highly distribut...
Jinxi Xu, W. Bruce Croft
SIGIR1
1998 Effective Retrieval with Distributed Collections
abstract
This paper evaluates the retrieval effectiveness of distributed information retrieval systems in realistic environments. We find that when a large number of collections are available, the retrieval effectiveness is significantly worse than that of centralized systems, mainly because typical queries are not adequate for the purpose of choosing the right collections. We propose two techniques to address the problem. One is to use phrase information in the collection selection index and the other is query expansion. Both techniques enhance the discriminatory power of typical queries for choosing the right collections and hence significantly improve retrieval results. Query expansion, in particular, brings the effectiveness of searching a large set of distributed collections close to that of searching a centralized collection. 1 Introduction In today's network environments, information is highly distributed. The Internet or World Wide Web, for example, contains thousands of collections. ...
Jinxi Xu, Jamie Callan
SIGIR1
1998 Corpus-Based Stemming Using Cooccurrence of Word Variants
abstract
Stemming is used in many information retrieval (IR) systems to reduce variant word forms to common roots. It is one of the simplest applications of natural-language processing to IR and is one of the most effective in terms of user acceptance and consistency, though small retrieval improvements. Current stemming techniques do not, however, reflect the language use in specific corpora, and this can lead to occasional serious retrieval failures. We propose a technique for using corpus-based word variant cooccurrence statistics to modify or create a stemmer. The experimental results generated using English newspaper and legal text and Spanish text demonstrate the viability of this technique and its advantages relative to conventional approaches that only employ morphological rules.
Jinxi Xu, W. Bruce Croft
ACM Trans. Inf. Syst.1
1996 Query Expansion Using Local and Global Document Analysis
abstract
Automatic query expansion has long been suggested as a technique for dealing with the fundamental issue of word mismatch in information retrieval.A number of approaches to ezpanrnion have been studied and, more recently, attention has focused on techniques that analyze the corpus to discover word relationship (global techniques) and those that analyze documents retrieved by the initial quer~( local feedback).In this paper, we compare the effectiveness of these approaches and show that, although global analysis haa some advantages, local analysia is generally more effective.We also show that using global analysis techniques, such as word contezt and phrase structure, on the local aet of documents produces results that are both more effective and more predictable than simple local feedback.
Jinxi Xu, W. Bruce Croft
SIGIR1