Xueqiang Lv

dblp:30/6106 · DBLP profile ↗
← Back
10ranked-venue papers in the field
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (2 first)Information Retrieval & Web Search · 3Database Systems & Data Management · 2
YearPublicationVenuePosition
2025 Sentence Extraction Framework with High Relevance and Divergence for Document Summarization
Huiwen Xue, Baoan Li, Denghao Ma, Xueqiang Lv, Xiaoxi Wang
DASFAA (1)4
2024 Reusing Keywords for Fine-grained Representations and Matchings
Li Chong, Denghao Ma, Yueguo Chen, Xueqiang Lv
DASFAA (2)4
2023 A Principled Decomposition of Pointwise Mutual Information for Intention Template Discovery
abstract
With the rise of Artificial Intelligence (AI), question answering systems have become common for users to interact with computers, e.g., ChatGPT and Siri. These systems require a substantial amount of labeled data to train their models. However, the labeled data is scarce and challenging to be constructed. The construction process typically involves two stages: discovering potential sample candidates and manually labeling these candidates. To discover high-quality candidate samples, we study the intention paraphrase template discovery task: Given some seed questions or templates of an intention, discover new paraphrase templates that describe the intention and are diverse to the seeds enough in text. As the first exploration of the task, we identify the new quality requirements, i.e., relevance, divergence and popularity, and identify the new challenges, i.e., the paradox of divergent yet relevant paraphrases, and the conflict of popular yet relevant paraphrases. To untangle the paradox of divergent yet relevant paraphrases, in which the traditional bag of words falls short, we develop usage-centric modeling, which represents a question/template/answer as a bag of usages that users engaged (e.g., up-votes), and uses a usage-flow graph to interrelate templates, questions and answers. To balance the conflict of popular yet relevant paraphrases, we propose a new and principled decomposition for the well-known Pointwise Mutual Information from the usage perspective (usage-PMI), and then develop a Bayesian inference framework over the usage-flow graph to estimate the usage-PMI. Extensive experiments over three large CQA corpora show strong performance advantage over the baselines adopted from paraphrase identification task. We release 885,000 paraphrase templates of high quality discovered by our proposed PMI decomposition model, and the data is available in site https://github.com/Para-Questions/Intention\_template\_discovery.
Denghao Ma, Kevin Chen-Chuan Chang, Yueguo Chen, Xueqiang Lv
CIKM4
2023 HBert: A Long Text Processing Method Based on BERT and Hierarchical Attention Mechanisms
abstract
With the emergence of a large-scale pre-training model based on the transformer model, the effect of all-natural language processing tasks has been pushed to a new level. However, due to the high complexity of the transformer's self-attention mechanism, these models have poor processing ability for long text. Aiming at solving this problem, a long text processing method named HBert based on Bert and hierarchical attention neural network is proposed. Firstly, the long text is divided into multiple sentences whose vectors are obtained through the word encoder composed of Bert and the word attention layer. And the article vector is obtained through the sentence encoder that is composed of transformer and sentence attention. Then the article vector is used to complete the subsequent tasks. The experimental results show that the proposed HBert method achieves good results in text classification and QA tasks. The F1 value is 95.7% in longer text classification tasks and 75.2% in QA tasks, which are better than the state-of-the-art model longformer.
Xueqiang Lv, Zhaonan Liu, Xindong You
Int. J. Semantic Web Inf. Syst.1
2023 Research on the Generation of Patented Technology Points in New Energy Based on Deep Learning
abstract
Effective extraction of patent technology points in new energy fields is profitable, which motivates technological innovation and facilitates patent transformation and application. However, since patent data exists the ununiform distribution of technology points information, long length of term, and long sentences, technology point extraction faces the dilemmas of poor readability and logic confusion. To mitigate these problems, the article proposes a method to generate patent technology points called IGPTP—a two-stage strategy, which fuses the advantage of extractive and generative ways. IGPTP utilizes the RoBERTa+CNN model to obtain the key sentences of text and takes the output as input of UNILM (unified pre-trained language model). Simultaneously, it takes a multi-strategies integration technique to enhance the quality of patent technology points by combining the copy mechanism and external knowledge guidance model. Substantial experimental results manifest that IGPTP outperforms the current mainstream models, which can generate more coherent and richer text.
Haixiang Yang, Xindong You, Xueqiang Lv
Int. J. Semantic Web Inf. Syst.3
2020 Distant Supervised Relation Extraction via DiSAN-2CNN on a Feature Level
abstract
At present, the mainstream distant supervised relation extraction methods existed problems: the coarse granularity for coding the context feature information; the difficulty in capturing the long-term dependency in the sentence, and the difficulty in coding prior knowledge of structures are major issues. To address these problems, we propose a distant supervised relation extraction model via DiSAN-2CNN on feature level, in which multi-dimension self-attention mechanism is utilized to encode the features of the words and DiSAN-2CNN is used to encode the sentence to obtain the long-term dependency, the prior knowledge of the structure, the time sequence, and the entity dependence in the sentence. Experiments conducted on the NYT-Freebase benchmark dataset demonstrate that the proposed DiSAN-2CNN on a feature level model achieves better performance than the current two state-of-art distant supervised relation extraction models PCNN+ATT and ResCNN-9, and it has d generalization ability with the least artificial feature engineering.
Xueqiang Lv, Huixin Hou, Xindong You, Junmei Han
Int. J. Semantic Web Inf. Syst.1
2015 Research on Semantic Disambiguation in Treebank
Xueqiang Lv, Yunfang Wu
APWeb2
2015 Location Prediction of Social Images via Generative Model
abstract
The vast amount of geo-tagged social images has attracted great attention in research of predicting location using the plentiful content of images, such as visual content and textual description. Most of the existing researches use the text-based or vision-based method to predict location. There still exists a problem: how to effectively exploit the correlation between different types of content as well as their geographical distributions for location prediction. In this paper, we propose to predict image location by learning the latent relation between geographical location and multiple types of image content. In particularly, we propose a geographical topic model GTMSI (geographical topic model of social image) to integrate multiple types of image content as well as the geographical distributions. In GTMI, image topic is modeled on both text vocabulary and visual feature. Each region has its own distribution over topics and hence has its own language model and vision pattern. The location of a new image is estimated based on the joint probability of image content and similarity measure on topic distribution between images. Experiment results demonstrate the performance of location prediction based on GTMSI.
Xiaoming Zhang 0001, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xueqiang Lv
ICMR5
2007 Design and Realization of Advertisement Promotion Based on the Content of Webpage
Shuicai Shi, Xueqiang Lv
KSEM3
2006 A Comparative Study on Representing Units in Chinese Text Clustering
Shiwen Yu, Xueqiang Lv, Shuicai Shi, Shibin Xiao
KSEM3