VLDB 2026 Research / reviewers in the wild / expert
Xiaoqiang Luo
dblp:10/4867
· DBLP profile ↗
21ranked-venue papers
11as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Information extraction and text analysis · 57% Machine translation · 34% Efficient and distributed learning · 7% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 86% Knowledge graphs · 14% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining
information extraction |
0.1 | 1 | 2012 | Distilling and exploring nuggets from a corpus · SIGIR 2012 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 3 | 2013 | Enlisting the Ghost: Modeling Empty Categories for Machine Translation · ACL (1) 2013 A Maximum Entropy Chinese Character-Based Parser · EMNLP 2003 Active Learning for Statistical Natural Language Parsing · ACL 2002 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.1 | 1 | 2011 | Learning to Transform and Select Elementary Trees for Improved Syntax-based Machine Translations · ACL 2011 |
Data mining
pattern mining |
0.1 | 1 | 2011 | A Statistical Tree Annotator and Its Applications · ACL 2011 |
Algorithms and data structures
tree data structures |
0.1 | 1 | 2011 | A Statistical Tree Annotator and Its Applications · ACL 2011 |
Natural language and speech › Information extraction and text analysis › named entity recognition
mention detection |
0.1 | 1 | 2009 | A Cascaded Approach to Mention Detection and Chaining in Arabic · IEEE Trans. Speech Audio Process. 2009 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.0 | 1 | 2004 | A Mention-Synchronous Coreference Resolution Algorithm Based On the Bell Tree · ACL 2004 |
Knowledge graphs › ontology
ontology matching |
0.0 | 1 | 2012 | Distilling and exploring nuggets from a corpus · SIGIR 2012 |
Natural language and speech › Information extraction and text analysis › named entity recognition
chinese named entity recognition |
0.0 | 1 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.0 | 1 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 |
Natural language and speech › Information extraction and text analysis
word segmentation |
0.0 | 1 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 |
Natural language and speech › Machine translation › syntax-based machine translation
tree-to-string translation |
0.0 | 1 | 2011 | Learning to Transform and Select Elementary Trees for Improved Syntax-based Machine Translations · ACL 2011 |
Machine learning › Efficient and distributed learning
active learning |
0.0 | 1 | 2002 | Active Learning for Statistical Natural Language Parsing · ACL 2002 |
Machine learning › Efficient and distributed learning › active learning
selective labeling |
0.0 | 1 | 2002 | Active Learning for Statistical Natural Language Parsing · ACL 2002 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing |
0.0 | 1 | 2002 | Active Learning for Statistical Natural Language Parsing · ACL 2002 |
Natural language and speech › Information extraction and text analysis › word segmentation
chinese word segmentation |
0.0 | 1 | 1996 | An Iterative Algorithm to Build Chinese Language Models · ACL 1996 |
Machine learning › Deep learning architectures and training
iterative training |
0.0 | 1 | 1996 | An Iterative Algorithm to Build Chinese Language Models · ACL 1996 |
Natural language and speech › Information extraction and text analysis › word segmentation
unsupervised word segmentation |
0.0 | 1 | 1996 | An Iterative Algorithm to Build Chinese Language Models · ACL 1996 |
Methods — techniques the papers use, named apart from their topics
statistical annotation · 0.2maximum entropy · 0.1elementary tree transformation · 0.1cascaded classification · 0.1maximum entropy model · 0.0bell tree · 0.0segmentation · 0.0combination · 0.0entropy-based uncertainty · 0.0clustering · 0.0active learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Enlisting the Ghost: Modeling Empty Categories for Machine Translation
Bing Xiang, Xiaoqiang Luo, Bowen Zhou 0006 |
ACL (1) | 2 |
| 2013 | Finding What Matters in Questions
Xiaoqiang Luo, Hema Raghavan, Vittorio Castelli, Sameer Maskey, Radu Florian |
HLT-NAACL | 1 |
| 2012 | Distilling and exploring nuggets from a corpusabstractThis paper describes a live and scalable system that automatically extracts information nuggets for entities/topics from a continuously updated corpus for effective exploration and analysis. A nugget is a piece of semantic information that (1) must be mapped semantically to the transitive closure of a pre-defined ontology, (2) is explicitly supported by text, and (3) has a natural language description that completely conveys its semantic to a user. Fig. 1 shows a type of nugget "involvement in events" for a person entity (Leon Panetta): each nugget has a short description ("meeting", "news conference") with a list of supporting passages. Vittorio Castelli, Hema Raghavan, Radu Florian, Ding-Jung Han, Xiaoqiang Luo, Salim Roukos |
SIGIR | 5 |
| 2011 | A Statistical Tree Annotator and Its Applications
Xiaoqiang Luo |
ACL | 1 |
| 2011 | Learning to Transform and Select Elementary Trees for Improved Syntax-based Machine Translations
Young-Suk Lee 0001, Xiaoqiang Luo |
ACL | 3 |
| 2010 | Learning to Predict Readability using Diverse Linguistic Features
Rohit J. Kate, Xiaoqiang Luo, Siddharth Patwardhan, Martin Franz, Radu Florian, Raymond J. Mooney, Salim Roukos, Christopher A. Welty |
COLING | 2 |
| 2010 | Constituent Reordering and Syntax Models for English-to-Japanese Statistical Machine Translation
Young-Suk Lee 0001, Xiaoqiang Luo |
COLING | 3 |
| 2009 | A Cascaded Approach to Mention Detection and Chaining in ArabicabstractThis paper presents a fully statistical approach to Arabic mention detection and chaining system, built around the maximum entropy principle. The presented system takes a cascade approach to processing an input document, by first detecting mentions in the document and then chaining the identified mentions into entities. Both system components use a common maximum entropy framework, which allows the integration of a large array of feature types, including lexical, morphological, syntactic, and semantic features. Arabic offers additional challenges for this task (when compared with English, for example), as segmentation is a needed processing step, so one can correctly identify and resolve enclitic pronouns. The system presented has obtained very competitive performance in the automatic content extraction (ACE) evaluation program. Imed Zitouni, Xiaoqiang Luo, Radu Florian |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Coreference or Not: A Twin Model for Coreference Resolution
Xiaoqiang Luo |
HLT-NAACL | 1 |
| 2004 | A Mention-Synchronous Coreference Resolution Algorithm Based On the Bell TreeabstractThis paper proposes a new approach for coreference resolution which uses the Bell tree to represent the search space and casts the coreference resolution problem as finding the best path from the root of the Bell tree to the leaf nodes. A Maximum Entropy model is used to rank these paths. The coreference performance on the 2002 and 2003 Automatic Content Extraction (ACE) data will be reported. We also train a coreference system using the MUC6 data and competitive results are obtained. Xiaoqiang Luo, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Salim Roukos |
ACL | 1 |
| 2004 | A Statistical Model for Multilingual Entity Detection and Tracking
Radu Florian, Hany Hassan, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Xiaoqiang Luo, Nicolas Nicolov, Salim Roukos |
HLT-NAACL | 6 |
| 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues
Hongyan Jing, Radu Florian, Xiaoqiang Luo, Tong Zhang 0001, Abraham Ittycheriah |
EMNLP | 3 |
| 2003 | A Maximum Entropy Chinese Character-Based Parser
Xiaoqiang Luo |
EMNLP | 1 |
| 2002 | Active Learning for Statistical Natural Language ParsingabstractIt is necessary to have a (large) annotated corpus to build a statistical parser. Acquisition of such a corpus is costly and time-consuming. This paper presents a method to reduce this demand using active learning, which selects what samples to annotate, instead of annotating blindly the whole training corpus.Sample selection for annotation is based upon "representativeness" and "usefulness". A model-based distance is proposed to measure the difference of two sentences and their most likely parse trees. Based on this distance, the active learning process analyzes the sample distribution by clustering and calculates the density of each sample to quantify its representativeness. Further more, a sentence is deemed as useful if the existing model is highly uncertain about its parses, where uncertainty is measured by various entropy-based scores.Experiments are carried out in the shallow semantic parser of an air travel dialog system. Our result shows that for about the same parsing accuracy, we only need to annotate a third of the samples as compared to the usual random selection method. Xiaoqiang Luo, Salim Roukos |
ACL | 2 |
| 2000 | Parser adaptation via Householder transformabstractWe propose a method of adapting a statistical parser using a special orthogonal transform, the Householder transform. Probability mass functions (pmf) in the parser are first mapped to unit sphere, then the Householder transform is applied, which maps a point in unit sphere to another point in unit sphere. The final model is obtained by mapping the transformed point in unit sphere back to simplex through a square map. The proposed method is tested on a semantic parser, and over 20% relative reduction of parse errors can be achieved. Xiaoqiang Luo |
ICASSP | 1 |
| 2000 | Semantic tokenization of verbalized numbers in language modeling
Xiaoqiang Luo, Martin Franz |
INTERSPEECH | 1 |
| 1999 | Probabilistic classification of HMM states for large vocabulary continuous speech recognitionabstractIn state-of-art large vocabulary continuous speech recognition (LVCSR) systems, HMM state-tying is often used to achieve good balance between the model resolution and robustness. In this paradigm, tied HMM states share a single set of parameters and are nondistinguishable. To capture the fine differences among tied HMM states, a probabilistic classification of HMM states (PCHMM) is proposed in this paper for LVCSR. In particular, a distribution from a HMM state to classes is introduced. It is shown that the state-to-class distribution can be estimated together with conventional HMM parameters within the EM (Dempster et al., 1977) framework. Compared with HMM state-tying, probabilistic classification of HMM states makes more efficient use of model parameters. It also makes the acoustic model more robust against the possible mismatch or variation between training and test data. The viability of this approach is verified by the significant reduction of word error rate (WER) on the Switchboard (Godfrey et al., 1992) task. Xiaoqiang Luo, Frederick Jelinek |
ICASSP | 1 |
| 1998 | Growth transform of a sum of rational functions and its application in estimating HMM parameters
Xiaoqiang Luo |
ICSLP | 1 |
| 1998 | Nonreciprocal data sharing in estimating HMM parameters
Xiaoqiang Luo, Frederick Jelinek |
ICSLP | 1 |
| 1998 | Speaker normalization with all-pass transformsabstractSpeaker normalization is a process in which the short-time features of speech from a given speaker are transformed so as to better match some speaker independent model. Vocal tract length normalization (VTLN) is a popular speaker normalization scheme wherein the frequency axis of the short-time spectrum associated with a particular speaker's speech is rescaled or warped prior to the extraction of cepstral features. In this work, we develop a novel speaker normalization scheme by exploiting the fact that frequency domain transformations similar to that inherent in VTLN can be accomplished entirely in the cepstral domain through the use of conformal maps. We propose a class of such maps, designated all-pass transforms for reasons given hereafter, and rigorously investigate their properties. Theoretical results are provided relating to the transformation of cepstral sequences under these maps. Additionally, all relations necessary to determine maximum likelihood estimates of the mapping parameters are derived for both speaker normalization and adaptation. 1 2 Speaker Normalization with All-Pass Transforms 1 John W. McDonough, William J. Byrne, Xiaoqiang Luo |
ICSLP | 3 |
| 1996 | An Iterative Algorithm to Build Chinese Language ModelsabstractWe present an iterative procedure to build a Chinese language model (LM). We segment Chinese text into words based on a word-based Chinese language model. However, the construction of a Chinese LM itself requires word boundaries. To get out of the chicken-and-egg problem, we propose an iterative procedure that alternates two operations: segmenting text into words and building an LM. Starting with an initial segmented corpus and an LM based upon it, we use a Viterbi-liek algorithm to segment another set of data. Then, we build an LM based on the second set and use the resulting LM to segment again the first corpus. The alternating procedure provides a self-organized way for the segmenter to detect automatically unseen words and correct segmentation errors. Our preliminary experiment shows that the alternating procedure not only improves the accuracy of our segmentation, but discovers unseen words suprisingly well. The resulting word-based LM has a perplexity of 188 for a general Chinese corpus. Xiaoqiang Luo, Salim Roukos |
ACL | 1 |