VLDB 2026 Research / reviewers in the wild / expert
Jinxi Xu
dblp:67/1760
· DBLP profile ↗
18ranked-venue papers
12as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 9 first-authorArtificial intelligence and machine learning · 7 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Information retrieval · 95% Data mining · 5% | |
| Artificial intelligence
5 papers |
Machine translation · 76% Language models and text generation · 12% Question answering and dialogue systems · 7% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
statistical machine translation |
0.3 | 3 | 2010 | Statistical Machine Translation with a Factorized Grammar · EMNLP 2010 Effective Use of Linguistic and Contextual Information for Statistical Machine Translation · EMNLP 2009 A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008 |
Information retrieval › query reformulation
query expansion |
0.3 | 5 | 2013 | Query expansion using path-constrained random walks · SIGIR 2013 Empirical studies in strategies for Arabic retrieval · SIGIR 2002 Improving the effectiveness of information retrieval with local context analysis · ACM Trans. Inf. Syst. 2000 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.1 | 2 | 2010 | A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008 Statistical Machine Translation with a Factorized Grammar · EMNLP 2010 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2010 | Combining Unsupervised and Supervised Alignments for MT: An Empirical Study · EMNLP 2010 |
Information retrieval
cross-language information retrieval |
0.1 | 3 | 2002 | Empirical studies in strategies for Arabic retrieval · SIGIR 2002 Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001 Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000 |
Natural language and speech › Language models and text generation › language modeling
dependency language model |
0.1 | 1 | 2008 | A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model · ACL 2008 |
Information retrieval
retrieval models |
0.1 | 3 | 2001 | Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001 Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999 Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000 |
Information retrieval
query log analysis |
0.0 | 1 | 2013 | Query expansion using path-constrained random walks · SIGIR 2013 |
Information retrieval
web search |
0.0 | 1 | 2013 | Query expansion using path-constrained random walks · SIGIR 2013 |
Natural language and speech › Question answering and dialogue systems › open-ended question answering
definitional question answering |
0.0 | 1 | 2004 | Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004 |
Data mining › text mining
information extraction and text analysis |
0.0 | 1 | 2004 | Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004 |
Information retrieval
distributed information retrieval |
0.0 | 2 | 1999 | Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999 Effective Retrieval with Distributed Collections · SIGIR 1998 |
Information retrieval › distributed information retrieval
resource selection |
0.0 | 2 | 1999 | Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999 Effective Retrieval with Distributed Collections · SIGIR 1998 |
Information retrieval › retrieval models
probabilistic retrieval model |
0.0 | 2 | 2001 | Evaluating a Probabilistic Model for Cross-Lingual Information Retrieval · SIGIR 2001 Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000 |
Information retrieval › cross-language information retrieval
arabic retrieval |
0.0 | 1 | 2002 | Empirical studies in strategies for Arabic retrieval · SIGIR 2002 |
Information retrieval › document retrieval › concept-based retrieval
thesaurus-based retrieval |
0.0 | 1 | 2002 | Empirical studies in strategies for Arabic retrieval · SIGIR 2002 |
Information retrieval › query reformulation › query expansion
local context analysis |
0.0 | 1 | 2000 | Improving the effectiveness of information retrieval with local context analysis · ACM Trans. Inf. Syst. 2000 |
Information retrieval › cross-language information retrieval
query translation |
0.0 | 1 | 2000 | Cross-lingual Information Retrieval Using Hidden Markov Models · EMNLP 2000 |
Information retrieval › retrieval models › language model
cluster-based language model |
0.0 | 1 | 1999 | Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999 |
Information retrieval › retrieval models
language model |
0.0 | 1 | 1999 | Cluster-Based Language Models for Distributed Retrieval · SIGIR 1999 |
Information retrieval › indexing
stemming |
0.0 | 1 | 1998 | Corpus-Based Stemming Using Cooccurrence of Word Variants · ACM Trans. Inf. Syst. 1998 |
Information retrieval
evaluation |
0.0 | 1 | 2004 | Evaluation of an extraction-based approach to answering definitional questions · SIGIR 2004 |
Methods — techniques the papers use, named apart from their topics
random walk · 0.2learning to rank · 0.2factorized grammar · 0.1empirical study · 0.1statistical machine translation · 0.1information extraction · 0.1ROUGE · 0.1target dependency language model · 0.1dependency parsing · 0.1stemming · 0.0spelling normalization · 0.0probabilistic term translation · 0.0character tri-grams · 0.0machine translation · 0.0generative model · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Query expansion using path-constrained random walksabstractThis paper exploits Web search logs for query expansion (QE) by presenting a new QE method based on path-constrained random walks (PCRW), where the search logs are represented as a labeled, directed graph, and the probability of picking an expansion term for an input query is computed by a learned combination of constrained random walks on the graph. The method is shown to be generic in that it covers most of the popular QE models as special cases and flexible in that it provides a principled mathematical framework in which a wide variety of information useful for QE can be incorporated in a unified way. Evaluation is performed on the Web document ranking task using a real-world data set. Results show that the PCRW-based method is very effective for the expansion of rare queries, i.e., low-frequency queries that are unseen in search logs, and that it outperforms significantly other state-of-the-art QE meth-ods. Jianfeng Gao 0001, Gu Xu, Jinxi Xu |
SIGIR | 3 |
| 2010 | Statistical Machine Translation with a Factorized Grammar
Libin Shen, Bing Zhang 0004, Spyridon Matsoukas, Jinxi Xu, Ralph M. Weischedel |
EMNLP | 4 |
| 2010 | Combining Unsupervised and Supervised Alignments for MT: An Empirical Study
Jinxi Xu, Antti-Veikko I. Rosti |
EMNLP | 1 |
| 2010 | String-to-Dependency Statistical Machine TranslationabstractWe propose a novel string-to-dependency algorithm for statistical machine translation. This algorithm employs a target dependency language model during decoding to exploit long distance word relations, which cannot be modeled with a traditional n-gram language model. Experiments show that the algorithm achieves significant improvement in MT performance over a state-of-the-art hierarchical string-to-string system on NIST MT06 and MT08 newswire evaluation sets. Libin Shen, Jinxi Xu, Ralph M. Weischedel |
Comput. Linguistics | 2 |
| 2009 | Effective Use of Linguistic and Contextual Information for Statistical Machine Translation
Libin Shen, Jinxi Xu, Bing Zhang 0004, Spyridon Matsoukas, Ralph M. Weischedel |
EMNLP | 2 |
| 2008 | A New String-to-Dependency Machine Translation Algorithm with a Target Dependency Language Model
Libin Shen, Jinxi Xu, Ralph M. Weischedel |
ACL | 2 |
| 2005 | Empirical studies on the impact of lexical resources on CLIR performance
Jinxi Xu, Ralph M. Weischedel |
Inf. Process. Manag. | 1 |
| 2004 | Evaluation of an extraction-based approach to answering definitional questionsabstractThis paper evaluates an extraction-based approach to answering definitional questions. Our system extracted useful linguistic constructs called linguistic features from raw text using information extraction tools and formulated answers based on such features. The features employed include appositives, copulas, structured patterns, relations, propositions and raw sentences. The features were ranked based on feature type and similarity to a question profile. Redundant features were detected using a simple heuristic-based strategy. The approach achieved state of the art performance at the TREC 2003 QA evaluation. Component analysis of the system was carried out using an automatic scoring function called Rouge (Lin and Hovy, 2003). Major findings include 1) answers using linguistic features are significantly better than those using raw sentences; 2) the most useful features are appositives and copulas; 3) question profiles, as a means of modeling user interests, can significantly improve system performance; 4) the Rouge scores are closely correlated with subjective evaluation results, indicating the suitability of using Rouge for evaluating definitional QA systems. Jinxi Xu, Ralph M. Weischedel, Ana Licuanan |
SIGIR | 1 |
| 2003 | Cross-lingual retrieval for HindiabstractIn this paper we describe the evaluation results of applying a cross-lingual retrieval model to retrieve Hindi documents relevant to an English query. Though the technique has been previously applied and evaluated for retrieving Chinese and Arabic documents given an English query, what is new about these experiments is porting the model to Hindi in two weeks' time. Jinxi Xu, Ralph M. Weischedel |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2002 | Audio Indexing of Arabic broadcast newsabstractThis paper describes the development of the BBN Audio Indexing System for broadcast news in Arabic. Key issues addressed in this work revolve around the three major components of the audio indexing system: automatic speech recognition, speaker identification, and named entity identification. The system deals with several challenges introduced by the Arabic language, including the absence of short vowels in written text and the presence of compound words that are formed by the concatenation of certain conjunctions, prepositions, articles, and pronouns, as prefixes and suffixes to the word stem. The lack of short vowels in the transcripts prompted a novel solution that further demonstrated the power of hidden Markov models to deal with ambiguity. Another challenge was the acquisition of appropriate language modeling data, given the absence of broadcast news data for that purpose. We present performance results for all three components of the Audio Indexing System, which we believe represent the state of the art for Arabic broadcast news. Jayadev Billa, Mohammed Noamany, Amit Srivastava, Daben Liu, Rebecca Stone, Jinxi Xu, John Makhoul, Francis Kubala |
ICASSP | 6 |
| 2002 | Empirical studies in strategies for Arabic retrievalabstractThis work evaluates a few search strategies for Arabic monolingual and cross-lingual retrieval, using the TREC Arabic corpus as the test-bed. The release by NIST in 2001 of an Arabic corpus of nearly 400k documents with both monolingual and cross-lingual queries and relevance judgments has been a new enabler for empirical studies. Experimental results show that spelling normalization and stemming can significantly improve Arabic monolingual retrieval. Character tri-grams from stems improved retrieval modestly on the test corpus, but the improvement is not statistically significant. To further improve retrieval, we propose a novel thesaurus-based technique. Different from existing approaches to thesaurus-based retrieval, ours formulates word synonyms as probabilistic term translations that can be automatically derived from a parallel corpus. Retrieval results show that the thesaurus can significantly improve Arabic monolingual retrieval. For cross-lingual retrieval (CLIR), we found that spelling normalization and stemming have little impact. Jinxi Xu, Alexander Fraser 0001, Ralph M. Weischedel |
SIGIR | 1 |
| 2001 | Evaluating a Probabilistic Model for Cross-Lingual Information RetrievalabstractThis work proposes and evaluates a probabilistic cross-lingual retrieval system. The system uses a generative model to estimate the probability that a document in one language is relevant, given a query in another language. An important component of the model is translation probabilities from terms in documents to terms in a query. Our approach is evaluated when 1) the only resource is a manually generated bilingual word list, 2) the only resource is a parallel corpus, and 3) both resources are combined in a mixture model. The combined resources produce about 90% of monolingual performance in retrieving Chinese documents. For Spanish the system achieves 85% of monolingual performance using only a pseudo-parallel Spanish-English corpus. Retrieval results are comparable with those of the structural query translation technique (Pirkola, 1998) when bilingual lexicons are used for query translation. When parallel texts in addition to conventional lexicons are used, it achieves better retrieval results but requires more computation than the structural query translation technique. It also produces slightly better results than using a machine translation system for CLIR, but the improvement over the MT system is not significant. Jinxi Xu, Ralph M. Weischedel |
SIGIR | 1 |
| 2000 | Cross-lingual Information Retrieval Using Hidden Markov ModelsabstractThis paper presents empirical results in cross-lingual information retrieval using English queries to access Chinese documents (TREC-5 and TREC-6) and Spanish documents (TREC-4). Since our interest is in languages where resources may be minimal, we use an integrated probabilistic model that requires only a bilingual dictionary as a resource. We explore how a combined probability model of term translation and retrieval can reduce the effect of translation ambiguity. In addition, we estimate an upper bound on performance, if translation ambiguity were a solved problem. We also measure performance as a function of bilingual dictionary size. Jinxi Xu, Ralph M. Weischedel |
EMNLP | 1 |
| 2000 | Improving the effectiveness of information retrieval with local context analysisabstractTechniques for automatic query expansion have been extensively studied in information research as a means of addressing the word mismatch between queries and documents. These techniques can be categorized as either global or local. While global techniques rely on analysis of a whole collection to discover word relationships, local techniques emphasize analysis of the top-ranked documents retrieved for a query. While local techniques have shown to be more effective that global techniques in general, existing local techniques are not robust and can seriously hurt retrieved when few of the retrieval documents are relevant. We propose a new technique, called local context analysis, which selects expansion terms based on cooccurrence with the query terms within the top-ranked documents. Experiments on a number of collections, both English and non-English, show that local context analysis offers more effective and consistent retrieval results. Jinxi Xu, W. Bruce Croft |
ACM Trans. Inf. Syst. | 1 |
| 1999 | Cluster-Based Language Models for Distributed RetrievalabstractEffective retrieval in a distributed environment is an important but difficult problem. Lack of effectiveness appears to have three causes. First, collection selection based on word histograms is not appropriate for heterogeneous collections. Second, relevant documents are scattered over many collections and searching a few collections misses many relevant documents. Third, most existing collection selection metrics lack sound theoretical justifications and hence may not be well tuned to the problem. We propose a new approach to distributed retrieval based on document clustering and language modeling. Document clustering is used to organize collections around topics. Language modeling is used to properly represent topics and effectively select the right topics for a query. Based on these ideas, three methods are proposed to suit different environments. We show that all three methods improve effectiveness of distributed retrieval. 1 Introduction Information has become highly distribut... Jinxi Xu, W. Bruce Croft |
SIGIR | 1 |
| 1998 | Effective Retrieval with Distributed CollectionsabstractThis paper evaluates the retrieval effectiveness of distributed information retrieval systems in realistic environments. We find that when a large number of collections are available, the retrieval effectiveness is significantly worse than that of centralized systems, mainly because typical queries are not adequate for the purpose of choosing the right collections. We propose two techniques to address the problem. One is to use phrase information in the collection selection index and the other is query expansion. Both techniques enhance the discriminatory power of typical queries for choosing the right collections and hence significantly improve retrieval results. Query expansion, in particular, brings the effectiveness of searching a large set of distributed collections close to that of searching a centralized collection. 1 Introduction In today's network environments, information is highly distributed. The Internet or World Wide Web, for example, contains thousands of collections. ... Jinxi Xu, Jamie Callan |
SIGIR | 1 |
| 1998 | Corpus-Based Stemming Using Cooccurrence of Word VariantsabstractStemming is used in many information retrieval (IR) systems to reduce variant word forms to common roots. It is one of the simplest applications of natural-language processing to IR and is one of the most effective in terms of user acceptance and consistency, though small retrieval improvements. Current stemming techniques do not, however, reflect the language use in specific corpora, and this can lead to occasional serious retrieval failures. We propose a technique for using corpus-based word variant cooccurrence statistics to modify or create a stemmer. The experimental results generated using English newspaper and legal text and Spanish text demonstrate the viability of this technique and its advantages relative to conventional approaches that only employ morphological rules. Jinxi Xu, W. Bruce Croft |
ACM Trans. Inf. Syst. | 1 |
| 1996 | Query Expansion Using Local and Global Document AnalysisabstractAutomatic query expansion has long been suggested as a technique for dealing with the fundamental issue of word mismatch in information retrieval.A number of approaches to ezpanrnion have been studied and, more recently, attention has focused on techniques that analyze the corpus to discover word relationship (global techniques) and those that analyze documents retrieved by the initial quer~( local feedback).In this paper, we compare the effectiveness of these approaches and show that, although global analysis haa some advantages, local analysia is generally more effective.We also show that using global analysis techniques, such as word contezt and phrase structure, on the local aet of documents produces results that are both more effective and more predictable than simple local feedback. Jinxi Xu, W. Bruce Croft |
SIGIR | 1 |