Jingfang Xu

dblp:95/8 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
5since 2021 · last 2022
0000-0003-3699-7116ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Systems, architecture and hardware · 2Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Question answering and dialogue systems · 35% Machine translation · 23% Language models and text generation · 10%
Databases, data mining, and information retrieval
7 papers
Information retrieval · 96% Machine learning and data management · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 61% Parallel and multicore computing · 30% Distributed systems · 9%

Topics — the 30 heaviest of 52, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
1.742021
ComQA: Compositional Question Answering via Hierarchical Graph Neural Networks · WWW 2021
A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020
ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion · AAAI 2020
Natural language and speech › Machine translation
neural machine translation
1.132021
Neural Machine Translation With Explicit Phrase Alignment · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Improving the Transformer Translation Model with Document-Level Context · EMNLP 2018
Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization · ACL (1) 2017
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.722019
A Compact and Language-Sensitive Multilingual Translation Method · ACL (1) 2019
Three Strategies to Improve One-to-Many Multilingual Translation · EMNLP 2018
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.722018
Commonsense Knowledge Aware Conversation Generation with Graph Attention · IJCAI 2018
Assigning Personality/Profile to a Chatting Machine for Coherent Conversation Generation · IJCAI 2018
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
compositional question answering
0.512021
ComQA: Compositional Question Answering via Hierarchical Graph Neural Networks · WWW 2021
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.512021
Neural Machine Translation With Explicit Phrase Alignment · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Machine learning › Transfer learning and domain adaptation
cross-task transfer
0.512021
Transfer Learning for Sequence Generation: from Single-source to Multi-source · ACL/IJCNLP (1) 2021
Machine learning › Deep learning architectures and training
encoder-decoder architecture
0.422019
Three Strategies to Improve One-to-Many Multilingual Translation · EMNLP 2018
A Compact and Language-Sensitive Multilingual Translation Method · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
evidence extraction
0.412020
A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020
Natural language and speech › Question answering and dialogue systems › question generation
neural question generation
0.412020
Neural Question Generation with Answer Pivot · AAAI 2020
Natural language and speech › Question answering and dialogue systems
question generation
0.412020
Neural Question Generation with Answer Pivot · AAAI 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.412020
A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence-to-sequence generation
0.412020
Modeling Voting for System Combination in Machine Translation · IJCAI 2020
Natural language and speech › Machine translation
system combination
0.412020
Modeling Voting for System Combination in Machine Translation · IJCAI 2020
Information retrieval
question answering
0.412020
ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion · AAAI 2020
Natural language and speech › Machine translation › computer-assisted translation
automatic post-editing
0.412019
Learning to Copy for Automatic Post-Editing · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › text generation › neural text generation
copy mechanism
0.412019
Learning to Copy for Automatic Post-Editing · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.412019
Document Gated Reader for Open-Domain Question Answering · SIGIR 2019
Natural language and speech › Information extraction and text analysis
stance detection
0.412019
Exploring Answer Stance Detection with Recurrent Conditional Attention · AAAI 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense knowledge
0.312018
Commonsense Knowledge Aware Conversation Generation with Graph Attention · IJCAI 2018
Natural language and speech › Machine translation
document-level machine translation
0.312018
Improving the Transformer Translation Model with Document-Level Context · EMNLP 2018
Natural language and speech › Question answering and dialogue systems › dialogue generation
knowledge-grounded dialogue generation
0.312018
Commonsense Knowledge Aware Conversation Generation with Graph Attention · IJCAI 2018
Natural language and speech › Language models and text generation
text generation
0.312018
Assigning Personality/Profile to a Chatting Machine for Coherent Conversation Generation · IJCAI 2018
High-performance computing
large-scale graph processing
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
Parallel and multicore computing
parallel graph algorithms
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
High-performance computing
scientific computing systems
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
posterior regularization
0.312017
Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization · ACL (1) 2017
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
prior knowledge integration
0.312017
Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization · ACL (1) 2017
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology-based query answering
query rewriting
0.312017
Cross-Lingual Information Retrieve in Sogou Search · SIGIR 2017
Information retrieval
cross-language information retrieval
0.312017
Cross-Lingual Information Retrieve in Sogou Search · SIGIR 2017

Methods — techniques the papers use, named apart from their topics

abstractive answer generation · 0.9BERT · 0.9attention mechanism · 0.8transfer learning · 0.5pre-training · 0.5phrase-based search space · 0.5hierarchical graph neural network · 0.5decoding algorithm · 0.5joint learning · 0.4answer pivot · 0.4result translation · 0.3query translation · 0.3n-gram representation · 0.1metric learning · 0.1locality-sensitive hashing · 0.1sequential dependency model · 0.1full dependency model · 0.1ranking SVM · 0.1
YearPublicationVenuePosition
2022 Automatic pediatric congenital heart disease classification based on heart sound signal
Jingjing Ye, Haomin Li 0001, Jingfang Xu, Jihua Zhu, Die Li, Qiang Shu
Artif. Intell. Medicine7
2021 Transfer Learning for Sequence Generation: from Single-source to Multi-source
abstract
Xuancheng Huang, Jingfang Xu, Maosong Sun, Yang Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xuancheng Huang, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
ACL/IJCNLP (1)2
2021 WG4Rec: Modeling Textual Content with Word Graph for News Recommendation
abstract
News recommendation plays an indispensable role in acquiring daily news for users. Previous studies make great efforts to model high-order feature interactions between users and items, where various neural models are applied (e.g., RNN, GNN). However, we find that seldom efforts are made to get better representations for news. Most previous methods simply adopt pre-trained word embeddings to represent news and also suffer from cold-start users.
Shaoyun Shi, Weizhi Ma, Zhen Wang 0040, Min Zhang 0006, Jingfang Xu, Yiqun Liu 0001, Shaoping Ma
CIKM6
2021 ComQA: Compositional Question Answering via Hierarchical Graph Neural Networks
abstract
With the development of deep learning techniques and large scale datasets, the question answering (QA) systems have been quickly improved, providing more accurate and satisfying answers. However, current QA systems either focus on the sentence-level answer, i.e., answer selection, or phrase-level answer, i.e., machine reading comprehension. How to produce compositional answers has not been throughout investigated. In compositional question answering, the systems should assemble several supporting evidence from the document to generate the final answer, which is more difficult than sentence-level or phrase-level QA. In this paper, we present a large-scale compositional question answering dataset containing more than 120k human-labeled questions. The answer in this dataset is composed of discontiguous sentences in the corresponding document. To tackle the ComQA problem, we proposed a hierarchical graph neural networks, which represent the document from the low-level word to the high-level sentence. We also devise a question selection and node selection task for pre-training. Our proposed model achieves a significant improvement over previous machine reading comprehension methods and pre-training methods. Codes, dataset can be found at https://github.com/benywon/ComQA.
Bingning Wang, Weipeng Chen, Jingfang Xu
WWW4
2021 Neural Machine Translation With Explicit Phrase Alignment
abstract
While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads to three problems: it is hard to (1) interpret the translation process, (2) impose lexical constraints, and (3) impose structural constraints. These problems not only increase the difficulty of designing new architectures for neural machine translation, but also limit its applications in practice. To alleviate these problems, we propose to introduce explicit phrase alignment into the translation process of arbitrary neural machine translation models. The key idea is to build a search space similar to that of phrase-based statistical machine translation for neural machine translation where phrase alignment is readily available. We design a new decoding algorithm that can easily impose lexical and structural constraints. Experiments show that our approach makes the translation process of neural machine translation more interpretable without sacrificing translation quality. In addition, our approach achieves significant improvements in lexically and structurally constrained translation tasks.
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Neural Question Generation with Answer Pivot
abstract
Neural question generation (NQG) is the task of generating questions from the given context with deep neural networks. Previous answer-aware NQG methods suffer from the problem that the generated answers are focusing on entity and most of the questions are trivial to be answered. The answer-agnostic NQG methods reduce the bias towards named entities and increasing the model's degrees of freedom, but sometimes result in generating unanswerable questions which are not valuable for the subsequent machine reading comprehension system. In this paper, we treat the answers as the hidden pivot for question generation and combine the question generation and answer selection process in a joint model. We achieve the state-of-the-art result on the SQuAD dataset according to automatic metric and human evaluation.
Bingning Wang, Ting Tao, Jingfang Xu
AAAI5
2020 ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion
abstract
This paper presents the ReCO, a human-curated Chinese Reading Comprehension dataset on Opinion. The questions in ReCO are opinion based queries issued to commercial search engine. The passages are provided by the crowdworkers who extract the support snippet from the retrieved documents. Finally, an abstractive yes/no/uncertain answer was given by the crowdworkers. The release of ReCO consists of 300k questions that to our knowledge is the largest in Chinese reading comprehension. A prominent characteristic of ReCO is that in addition to the original context paragraph, we also provided the support evidence that could be directly used to answer the question. Quality analysis demonstrates the challenge of ReCO that it requires various types of reasoning skills such as causal inference, logical reasoning, etc. Current QA models that perform very well on many question answering problems, such as BERT (Devlin et al. 2018), only achieves 77% accuracy on this dataset, a large margin behind humans nearly 92% performance, indicating ReCO present a good challenge for machine reading comprehension. The codes, dataset and leaderboard will be freely available at https://github.com/benywon/ReCO.
Bingning Wang, Jingfang Xu
AAAI4
2020 A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction
abstract
Neural models have achieved great success on machine reading comprehension (MRC), many of which typically consist of two components: an evidence extractor and an answer predictor.The former seeks the most relevant information from a reference text, while the latter is to locate or generate answers from the extracted evidence.Despite the importance of evidence labels for training the evidence extractor, they are not cheaply accessible, particularly in many non-extractive MRC tasks such as YES/NO question answering and multi-choice MRC.To address this problem, we present a Self-Training method (STM), which supervises the evidence extractor with auto-generated evidence labels in an iterative process.At each iteration, a base MRC model is trained with golden answers and noisy evidence labels.The trained model will predict pseudo evidence labels as extra supervision in the next iteration.We evaluate STM on seven datasets over three MRC tasks.Experimental results demonstrate the improvement on existing MRC models, and we also analyze how and why such a self-training method works in MRC.
Yilin Niu, Fangkai Jiao, Mantong Zhou, Jingfang Xu, Minlie Huang
ACL5
2020 Modeling Voting for System Combination in Machine Translation
abstract
System combination is an important technique for combining the hypotheses of different machine translation systems to improve translation performance. Although early statistical approaches to system combination have been proven effective in analyzing the consensus between hypotheses, they suffer from the error propagation problem due to the use of pipelines. While this problem has been alleviated by end-to-end training of multi-source sequence-to-sequence models recently, these neural models do not explicitly analyze the relations between hypotheses and fail to capture their agreement because the attention to a word in a hypothesis is calculated independently, ignoring the fact that the word might occur in multiple hypotheses. In this work, we propose an approach to modeling voting for system combination in machine translation. The basic idea is to enable words in hypotheses from different systems to vote on words that are representative and should get involved in the generation process. This can be done by quantifying the influence of each voter and its preference for each candidate. Our approach combines the advantages of statistical and neural methods since it can not only analyze the relations between hypotheses but also allow for end-to-end training. Experiments show that our approach is capable of better taking advantage of the consensus between hypotheses and achieves significant improvements over state-of-the-art baselines on Chinese-English and English-German machine translation tasks.
Xuancheng Huang, Zhixing Tan, Derek F. Wong, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
IJCAI6
2020 Incorporating Knowledge and Content Information to Boost News Recommendation
Zhen Wang 0040, Weizhi Ma, Min Zhang 0006, Weipeng Chen, Jingfang Xu, Yiqun Liu 0001, Shaoping Ma
NLPCC (1)5
2019 Exploring Answer Stance Detection with Recurrent Conditional Attention
abstract
Detecting stance from certain types of question-answer pairs is an interesting problem which has not been carefully explored. Unlike previous stance detection tasks, targets here are not given entities or claims but entire questions, which makes it difficult to capture the semantics of targets and build target-dependent representations of answers. To address them, we introduce the Recurrent Conditional Attention (RCA) model which incorporates a conditional attention structure into the recurrent reading process. RCA iteratively guides the distillation of question semantic with answer information and collects stance-oriented text relating to question, further revealing mutual relationship among stance, answer and question. Experiments on a manually labeled Chinese community QA stance dataset show that RCA outperforms four strong baselines by average 2.90% on macro-F1 and 2.66% on micro-F1 respectively.
Jianhua Yuan, Jingfang Xu, Bing Qin 0001
AAAI3
2019 A Compact and Language-Sensitive Multilingual Translation Method
abstract
Multilingual neural machine translation (Multi-NMT) with one encoder-decoder model has made remarkable progress due to its simple deployment.However, this multilingual translation paradigm does not make full use of language commonality and parameter sharing between encoder and decoder.Furthermore, this kind of paradigm cannot outperform the individual models trained on bilingual corpus in most cases.In this paper, we propose a compact and language-sensitive method for multilingual translation.To maximize parameter sharing, we first present a universal representor to replace both encoder and decoder models.To make the representor sensitive for specific languages, we further introduce language-sensitive embedding, attention, and discriminator with the ability to enhance model performance.We verify our methods on various translation scenarios, including one-to-many, many-to-many and zero-shot.Extensive experiments demonstrate that our proposed methods remarkably outperform strong standard multilingual translation systems on WMT and IWSLT datasets.Moreover, we find that our model is especially helpful in low-resource and zero-shot translation scenarios.
Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong
ACL (1)5
2019 Learning to Copy for Automatic Post-Editing
abstract
Xuancheng Huang, Yang Liu, Huanbo Luan, Jingfang Xu, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xuancheng Huang, Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
EMNLP/IJCNLP (1)4
2019 Document Gated Reader for Open-Domain Question Answering
abstract
Open-domain question answering focuses on using diverse information resources to answer any types of question. Recent years, with the development of large-scale data set and various deep neural networks models, some recent advances in open domain question answering system first utilize the distantly supervised dataset as the knowledge resource, then apply deep learning based machine comprehension techniques to generate the right answers, which achieves impressive results compared with traditional feature-based pipeline methods.
Bingning Wang, Jingfang Xu, Zhixing Tian, Kang Liu 0001, Jun Zhao 0001
SIGIR4
2019 Privacy-preserving raw data collection without a trusted authority for IoT
Yi-Ning Liu 0002, Yan-Ping Wang, Xiao-Fen Wang, Zhe Xia, Jingfang Xu
Comput. Networks5
2018 Three Strategies to Improve One-to-Many Multilingual Translation
abstract
Due to the benefits of model compactness, multilingual translation (including many-toone, many-to-many and one-to-many) based on a universal encoder-decoder architecture attracts more and more attention.However, previous studies show that one-to-many translation based on this framework cannot perform on par with the individually trained models.In this work, we introduce three strategies to improve one-to-many multilingual translation by balancing the shared and unique features.Within the architecture of one decoder for all target languages, we first exploit the use of unique initial states for different target languages.Then, we employ language-dependent positional embeddings.Finally and especially, we propose to divide the hidden cells of the decoder into shared and language-dependent ones.The extensive experiments demonstrate that our proposed methods can obtain remarkable improvements over the strong baselines.Moreover, our strategies can achieve comparable or even better performance than the individually trained translation models.
Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong
EMNLP4
2018 Improving the Transformer Translation Model with Document-Level Context
abstract
Although the Transformer translation model (Vaswani et al., 2017) has achieved state-ofthe-art performance in a variety of translation tasks, how to use document-level context to deal with discourse phenomena problematic for Transformer still remains a challenge.In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder.As large-scale document-level parallel corpora are usually not available, we introduce a two-step training method to take full advantage of abundant sentence-level parallel corpora and limited document-level parallel corpora.Experiments on the NIST Chinese-English datasets and the IWSLT French-English datasets show that our approach improves over Transformer significantly. 1
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Min Zhang 0005, Yang Liu 0005
EMNLP5
2018 Assigning Personality/Profile to a Chatting Machine for Coherent Conversation Generation
abstract
Endowing a chatbot with personality is challenging but significant to deliver more realistic and natural conversations. In this paper, we address the issue of generating responses that are coherent to a pre-specified personality or profile. We present a method that uses generic conversation data from social media (without speaker identities) to generate profile-coherent responses. The central idea is to detect whether a profile should be used when responding to a user post (by a profile detector), and if necessary, select a key-value pair from the profile to generate a response forward and backward (by a bidirectional decoder) so that a personality-coherent response can be generated. Furthermore, in order to train the bidirectional decoder with generic dialogue data, a position detector is designed to predict a word position from which decoding should start given a profile value. Manual and automatic evaluation shows that our model can deliver more coherent, natural, and diversified responses.
Qiao Qian, Minlie Huang, Haizhou Zhao, Jingfang Xu, Xiaoyan Zhu 0001
IJCAI4
2018 Commonsense Knowledge Aware Conversation Generation with Graph Attention
abstract
Commonsense knowledge is vital to many natural language processing tasks. In this paper, we present a novel open-domain conversation generation model to demonstrate how large-scale commonsense knowledge can facilitate language understanding and generation. Given a user post, the model retrieves relevant knowledge graphs from a knowledge base and then encodes the graphs with a static graph attention mechanism, which augments the semantic information of the post and thus supports better understanding of the post. Then, during word generation, the model attentively reads the retrieved knowledge graphs and the knowledge triples within each graph to facilitate better generation through a dynamic graph attention mechanism. This is the first attempt that uses large-scale commonsense knowledge in conversation generation. Furthermore, unlike existing models that use knowledge triples (entities) separately and independently, our model treats each knowledge graph as a whole, which encodes more structured, connected semantic information in the graphs. Experiments show that the proposed model can generate more appropriate and informative responses than state-of-the-art baselines.
Hao Zhou 0012, Tom Young, Minlie Huang, Haizhou Zhao, Jingfang Xu, Xiaoyan Zhu 0001
IJCAI5
2018 Privacy-Preserving Data Collection for Mobile Phone Sensing Tasks
Yi-Ning Liu 0002, Yan-Ping Wang, Xiao-Fen Wang, Zhe Xia, Jingfang Xu
ISPEC5
2018 ShenTu: processing multi-trillion edge graphs on millions of cores in seconds
Heng Lin, Xiaowei Zhu 0001, Bowen Yu 0003, Xiongchao Tang, Wei Xue 0003, Lufei Zhang, Torsten Hoefler, Xiaosong Ma, Xin Liu 0081, Jingfang Xu
SC12
2017 Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization
abstract
Although neural machine translation has made significant progress recently, how to integrate multiple overlapping, arbitrary prior knowledge sources remains a challenge.In this work, we propose to use posterior regularization to provide a general framework for integrating prior knowledge into neural machine translation.We represent prior knowledge sources as features in a log-linear model, which guides the learning process of the neural translation model.Experiments on Chinese-English translation show that our approach leads to significant improvements.
Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
ACL (1)4
2017 DeepRank: A New Deep Architecture for Relevance Ranking in Information Retrieval
abstract
This paper concerns a deep learning approach to relevance ranking in information retrieval (IR). Existing deep IR models such as DSSM and CDSSM directly apply neural networks to generate ranking scores, without explicit understandings of the relevance. According to the human judgement process, a relevance label is generated by the following three steps: 1) relevant locations are detected; 2) local relevances are determined; 3) local relevances are aggregated to output the relevance label. In this paper we propose a new deep learning architecture, namely DeepRank, to simulate the above human judgment process. Firstly, a detection strategy is designed to extract the relevant contexts. Then, a measure network is applied to determine the local relevances by utilizing a convolutional neural network (CNN) or two-dimensional gated recurrent units (2D-GRU). Finally, an aggregation network with sequential integration and term gating mechanism is used to produce a global relevance score. DeepRank well captures important IR characteristics, including exact/semantic matching signals, proximity heuristics, query term importance, and diverse relevance requirement. Experiments on both benchmark LETOR dataset and a large scale clickthrough data show that DeepRank can significantly outperform learning to ranking methods, and existing deep learning methods.
Liang Pang 0001, Yanyan Lan, Jiafeng Guo, Jun Xu 0001, Jingfang Xu, Xueqi Cheng 0001
CIKM5
2017 SogouT-16: A New Web Corpus to Embrace IR Research
abstract
Web collection is essential for many Web based researches such as Web Information Retrieval (IR), Web data mining, Corpus linguistics and so on. However, it is usually expensive and time-consuming to collect a large scale of Web pages in lab-based environment and public-available collection becomes a necessity for these researches. In this study, we present a Chinese Web collection, SogouT-16, which is the largest free-of-charge public Chinese Web collection so far. We provide a variety of descriptive characteristics of SogouT-16 and discuss its adoption in a newly-designed ad-hoc retrieval task in NTCIR-13, We Want Web. SogouT-16 also provides online retrieval service and contains a number of auxiliary resources including hyperlink structure graph, query logs, word embedding, and etc. We believe that SogouT-16 will provide new opportunities for novel investigations and applications in IR and other related communities.
Cheng Luo 0001, Yukun Zheng, Yiqun Liu 0001, Jingfang Xu, Min Zhang 0006, Shaoping Ma
SIGIR5
2017 Cross-Lingual Information Retrieve in Sogou Search
abstract
In recent years, more and more Chinese people desires to be able to access the large amount of foreign language information and understand what is happening all over the world. However, language barrier is always a problem to them. In order to break the language barrier and connect Chinese people to the foreign language information in the world, Sogou has built a cross-lingual information retrieval (CLIR) system named Sogou English (http://english.sogou.com), which enables Chinese people to search and browse foreign language information with Chinese. In Sogou English, when the user inputs a Chinese query, it will first translate the Chinese query into English, and then search over the Internet, and finally translate the search results into Chinese so that users can understand them better. Hence with Sogou English, people can read and browse the information from English world without actually knowing English.
Jingfang Xu, Feifei Zhai, Zhengshan Xue
SIGIR1
2011 Data Selection for User Topic Model in Twitter-Like Service
abstract
Twitter-like services are now a popular kind of online social networking services, in which user can express themselves, share contents, and follow others they are interested in. User modeling, building a model for user's interests, is a key problem in many social networking applications, such as recommendation, advertisement, etc. This paper focuses on data selection for user modeling in Twitter-like services. That is, we study the problem of how to select useful data to model a user's interests. Using different data, three user modeling methods are proposed and experiments on a real Twitter-like service are conducted to verify the effectiveness of proposed approaches. Experimental results shows that modeling user's interests with what he/she wrote and selectively what he read performs the best among the three methods we proposed.
Jingfang Xu, Xing Li 0001
ICPADS2
2011 Learning similarity function for rare queries
abstract
The key element of many query processing tasks can be formalized as calculation of similarities between queries. These include query suggestion, query reformulation, and query expansion. Although many methods have been proposed for query similarity calculation, they could perform poorly on rare queries. As far as we know, there was no previous work particularly about rare query similarity calculation, and this paper tries to study this problem. Specifically, we address three problems. Firstly, we define an n-gram space to represent queries with their own content and a similarity function to measure the similarities between queries. Secondly, we propose learning the similarity function by leveraging the training data derived from user behavior data. This is formalized as an optimization problem and a metric learning approach is employed to solve it efficiently. Finally, we exploit locality sensitive hashing for efficient retrieval of similar queries from a large query repository. We experimentally verified the effectiveness of the proposed approach by showing that our method can indeed enhance the accuracy of query similarity calculation for rare queries and efficiently retrieve similar queries. As an application, we also experimentally demonstrated that the similar queries found by our method can significantly improve search relevance.
Jingfang Xu, Gu Xu
WSDM1
2010 Improving quality of training data for learning to rank using click-through data
abstract
In information retrieval, relevance of documents with respect to queries is usually judged by humans, and used in evaluation and/or learning of ranking functions. Previous work has shown that certain level of noise in relevance judgments has little effect on evaluation, especially for comparison purposes. Recently learning to rank has become one of the major means to create ranking models in which the models are automatically learned from the data derived from a large number of relevance judgments. As far as we know, there was no previous work about quality of training data for learning to rank, and this paper tries to study the issue. Specifically, we address three problems. Firstly, we show that the quality of training data labeled by humans has critical impact on the performance of learning to rank algorithms. Secondly, we propose detecting relevance judgment errors using click-through data accumulated at a search engine. Two discriminative models, referred to as sequential dependency model and full dependency model, are proposed to make the detection. Both models consider the conditional dependency of relevance labels and thus are more powerful than the conditionally independent model previously proposed for other tasks. Finally, we verify that using training data in which the errors are detected and corrected by our method, we can improve the performance of learning to rank algorithms.
Jingfang Xu, Chuanliang Chen, Gu Xu, Hang Li 0001, Elbio Renato Torres Abib
WSDM1
2007 Learning to rank collections
abstract
Collection selection, ranking collections according to user query is crucial in distributed search. However, few features are used to rank collections in the current collection selection methods, while hundreds of features are exploited to rank web pages in web search. The lack of features affects the efficiency of collection selection in distributed search. In this paper, we exploit some new features and learn to rank collections with them through SVM and RankingSVM respectively. Experimental results show that our features are beneficial to collection selection, and the learned ranking functions outperform the classical CORI algorithm.
Jingfang Xu, Xing Li 0001
SIGIR1
2007 Estimating collection size with logistic regression
abstract
Collection size is an important feature to represent the content summaries of a collection, and plays a vital role in collection selection for distributed search. In uncooperative environments, collection size estimation algorithms are adopted to estimate the sizes of collections with their search interfaces. This paper proposes heterogeneous capture (HC) algorithm, in which the capture probabilities of documents are modeled with logistic regression. With heterogeneous capture probabilities, HC algorithm estimates collection size through conditional maximum likelihood. Experimental results on real web data show that our HC algorithm outperforms both multiple capture-recapture and capture history algorithms.
Jingfang Xu, Sheng Wu 0002, Xing Li 0001
SIGIR1
2005 An Article Language Model for BBS Search
Jingfang Xu, Yangbo Zhu, Xing Li 0001
ICWE1
2004 Query Based Chinese Phrase Extraction for Site Search
Jingfang Xu, Shaozhi Ye, Xing Li 0001
WISE1