Xueqiang Lv

dblp:30/6106 · DBLP profile ↗
← Back
43ranked-venue papers
4as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 9 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 YOLO-IOD: Towards Real Time Incremental Object Detection
abstract
Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to catastrophic forgetting in YOLO-based incremental detectors: foreground-background confusion, parameter interference, and misaligned knowledge distillation. Subsequently, we introduce YOLO-IOD, a real-time Incremental Object Detection (IOD) framework that is constructed upon the pretrained YOLO-World model, facilitating incremental learning via a stage-wise parameter-efficient finetuning process. Specifically, YOLO-IOD encompasses three principal components: 1) Conflict-Aware Pseudo-Label Refinement (CPR), which mitigates the foreground-background confusion by leveraging the confidence levels of pseudo labels and identifying potential objects relevant to future tasks. 2) Importance-based Kernel Selection (IKS), which identifies and updates the pivotal convolution kernels pertinent to the current task during the current learning stage. 3)Cross-Stage Asymmetric Knowledge Distillation (CAKD), which addresses the misaligned knowledge distillation conflict by transmitting the features of the student target detector through the detection heads of both the previous and current teacher detectors, thereby facilitating asymmetric distillation between existing and newly introduced categories. We further introduce LoCo COCO, a more realistic benchmark that eliminates data leakage across stages. Experiments on both conventional and LoCo COCO benchmarks show that YOLO-IOD achieves superior performance with minimal forgetting.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Chen Zhao 0009, Yanning Zhang 0001
AAAI2
2026 Small Object Detection via Frequency-Based Multi-modal Fusion
Shangzhi Teng, Yekai Li, Xi Gong, Xueqiang Lv
MMM (1)4
2026 Semantic-driven seasonal data classification: An artificial intelligence-enabled cost-effective storage system
Xueqiang Lv, Yunchao Gong, Xiao Qin 0001, Xindong You
Eng. Appl. Artif. Intell.2
2025 Revisiting Generative Replay for Class Incremental Object Detection
abstract
Generative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation that the forgetting of prior knowledge is predominantly present in the classification sub-task as opposed to the localization sub-task, we revisit the generative replay method for class incremental object detection. Our method utilize a standard Stable Diffusion model to generate image-level replay data for all old and new tasks. Accordingly, the old detector and a stage-wise detector are conducted on the synthetic images respectively to determine the bounding box positions through pseudo-labeling. Furthermore, we propose to use a Similarity-based Cross Sampling mechanism to select valuable confusing data between old and new tasks to more effectively mitigate catastrophic forgetting and reduce the false alarm rate for the new task. Finally, all synthetic and real data are integrated for current-stage detector training, where the images generated for previous tasks are highly beneficial in minimizing the forgetting of existing knowledge, while those synthesized for the new task can help bridge the domain gap between real and synthetic images. We conducted extensive experiments on PASCAL VOC 2007 and MS COCO benchmark datasets in multiple settings to showcase the efficacy of our proposed approach, which achieves state-of-the-art results. The code is available at https://github.com/qiangzailv/RGR-IOD.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Yanning Zhang 0001
CVPR2
2025 Sentence Extraction Framework with High Relevance and Divergence for Document Summarization
Huiwen Xue, Baoan Li, Denghao Ma, Xueqiang Lv, Xiaoxi Wang
DASFAA (1)4
2025 CE-DCVSI: Multimodal relational extraction based on collaborative enhancement of dual-channel visual semantic information
Yunchao Gong, Xueqiang Lv, Zangtai Cai, Yuzhong Chen 0003, Zhaojun Wang, Xindong You
Expert Syst. Appl.2
2025 An unsupervised fusion framework of generation and retrieval for entity search
Denghao Ma, Xueqiang Lv, Yanhe Du, Changyu Wang, Hongbin Pei
Expert Syst. Appl.3
2025 Stronger interaction brings better performance: fine-grained alignment between different modalities for MNER
Xindong You, Jiang Jian, Xueqiang Lv
J. Supercomput.4
2024 Reusing Keywords for Fine-grained Representations and Matchings
Li Chong, Denghao Ma, Yueguo Chen, Xueqiang Lv
DASFAA (2)4
2024 GNN-Based Multimodal Named Entity Recognition
abstract
Abstract The Multimodal Named Entity Recognition (MNER) task enhances the text representations and improves the accuracy and robustness of named entity recognition by leveraging visual information from images. However, previous methods have two limitations: (i) the semantic mismatch between text and image modalities makes it challenging to establish accurate internal connections between words and visual representations. Besides, the limited number of characters in social media posts leads to semantic and contextual ambiguity, further exacerbating the semantic mismatch between modalities. (ii) Existing methods employ cross-modal attention mechanisms to facilitate interaction and fusion between different modalities, overlooking fine-grained correspondences between semantic units of text and images. To alleviate these issues, we propose a graph neural network approach for MNER (GNN-MNER), which promotes fine-grained alignment and interaction between semantic units of different modalities. Specifically, to mitigate the issue of semantic mismatch between modalities, we construct corresponding graph structures for text and images, and leverage graph convolutional networks to augment text and visual representations. For the second issue, we propose a multimodal interaction graph to explicitly represent the fine-grained semantic correspondences between text and visual objects. Based on this graph, we implement deep-level feature fusion between modalities utilizing graph attention networks. Compared with existing methods, our approach is the first to extend graph deep learning throughout the MNER task. Extensive experiments on the Twitter multimodal datasets validate the effectiveness of our GNN-MNER.
Yunchao Gong, Xueqiang Lv, Xindong You, Yuzhong Chen 0003
Comput. J.2
2024 CSEA: A Fine-Grained Framework of Climate-Season-Based Energy-Aware in Cloud Storage Systems
abstract
Abstract Continuous data scale growth increases energy consumption and operating cost that cannot be ignored in cloud storage systems. Previous studies have shown that analyzing the characteristics of I/O access and mining data features is effective for reasonable data distribution in storage systems. The granularity and criterion of classification are the key factors in determining the data distribution. To decrease energy consumption and operating cost, this paper puts forward a fine-grained framework of the climatic-season-based energy-aware in cloud storage system called CSEA. The framework concludes the following three aspects: (i) data feature mining. CSEA discovers potential data features by analyzing data access to provide help with data classification. (ii) K-means clustering algorithm. CSEA uses an unsupervised data classification algorithm in machine learning to divide data into categories based on seasonal characteristics by gathering real I/O access. (iii) data distribution of fine-grained. On the basis of seasonal features, CSEA fuses regional features to further refine the data distribution granularity to save on energy consumption and operating cost. Simulation experiments using extended CloudSimDisk and the constructed mathematical models indicate that CSEA reduces the energy consumption and operating cost compared with the single data classification standard and coarse-grained data distribution.
Xueqiang Lv, Haojie Ge, Xindong You
Comput. J.2
2024 Deep click interest network for reranking hotels
Denghao Ma, Hongbin Pei, Xueqiang Lv, Genliang Yi, Haoxing Wen
Eng. Appl. Artif. Intell.4
2024 Data-driven smoothing approaches for interest modeling in recommendation systems
Denghao Ma, Xiayu Wang, Xueqiang Lv, Hongbin Pei, Youyou Zhang
Expert Syst. Appl.3
2024 Cost-effective data classification storage through text seasonal features
abstract
Data classification storage has emerged as an effective strategy, harnessing the diverse performance attributes of storage devices and orchestrating a harmonious equilibrium between energy consumption, cost considerations, and user accessibility. As research on emerging storage media (e.g. Non-Volatile Memory (NVM)) delves deeper, in scenarios characterized by dynamically evolving storage demands, conventional paradigms of rigid classification strategy and Solid-State Drive (SSD) and Hard Disk Drive (HDD) storage architectures fall short of addressing such complex situations. In this paper, we propose an effective data classification storage method using text seasonal features based on the traditional access frequency analysis strategy. First, to procure richer semantic information, we employ external knowledge actualizing the short-text feature expansion. Then, we leverage the ensemble learning stacking method optimized models to improve the accuracy of data classification based on seasonal features. Additionally, following the seasonal feature, we further classify the data into hot, warm, and cold and place them separately on the NVM, SSD, and HDD, which conserves storage energy consumption and operational costs while ensuring the quality of user access. The experimental results demonstrate that data classification accuracy can reach more than 95.10%, and the energy consumption and operating cost can be reduced by more than 30.22% and 8.73%, respectively.
Xueqiang Lv, Yunchao Gong, Taifu Yuan, Xindong You
Future Gener. Comput. Syst.2
2024 RPEPL: Tibetan Sentiment Analysis Based on Relative Position Encoding and Prompt Learning
abstract
Sentiment analysis is a critical task for natural language processing. Much research has been done for high-resource languages such as English and Chinese. However, Tibetan is an extremely low resource language with less reference information. According to the practical demands, this article proposes RPEPL, a Tibetan sentiment analysis method based on relative position encoding and prompt learning. First, word information is introduced to syllable sequences by converting the directed acyclic lattice into a squashed structure. Second, a relative position encoding is used to encode the position information of syllables and words. Third, the association relations and semantic information of tokens are identified by leveraging the multi-attention. Finally, the sentiment category of the Tibetan sentence is obtained through a prompt learning framework. Experimental results demonstrate that RPEPL significantly outperforms the baseline methods on the TUSA dataset and TNEC (Tibetan News Event Comments) dataset. Additionally, traditional recurrent neural networks cannot perform large-scale parallel computation and convolutional neural networks have difficulty modeling long-distance dependencies in Tibetan text, both of which are resolved using RPEPL. Furthermore, the use of multi-attention not only enriches the association relations between syllables and words but also enhances the understanding of sentence semantic and syntactic structure information, and improves the performance of Tibetan sentiment analysis.
Chunwei Kong, Xueqiang Lv, Haixing Zhao, Zangtai Cai, Yuzhong Chen 0003
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Multimodal heterogeneous graph entity-level fusion for named entity recognition with multi-granularity visual guidance
Yunchao Gong, Xueqiang Lv, Zhaojun Wang, Xindong You
J. Supercomput.2
2024 A relation enhanced model for temporal knowledge graph alignment
Zhaojun Wang, Xindong You, Xueqiang Lv
J. Supercomput.3
2023 A Principled Decomposition of Pointwise Mutual Information for Intention Template Discovery
abstract
With the rise of Artificial Intelligence (AI), question answering systems have become common for users to interact with computers, e.g., ChatGPT and Siri. These systems require a substantial amount of labeled data to train their models. However, the labeled data is scarce and challenging to be constructed. The construction process typically involves two stages: discovering potential sample candidates and manually labeling these candidates. To discover high-quality candidate samples, we study the intention paraphrase template discovery task: Given some seed questions or templates of an intention, discover new paraphrase templates that describe the intention and are diverse to the seeds enough in text. As the first exploration of the task, we identify the new quality requirements, i.e., relevance, divergence and popularity, and identify the new challenges, i.e., the paradox of divergent yet relevant paraphrases, and the conflict of popular yet relevant paraphrases. To untangle the paradox of divergent yet relevant paraphrases, in which the traditional bag of words falls short, we develop usage-centric modeling, which represents a question/template/answer as a bag of usages that users engaged (e.g., up-votes), and uses a usage-flow graph to interrelate templates, questions and answers. To balance the conflict of popular yet relevant paraphrases, we propose a new and principled decomposition for the well-known Pointwise Mutual Information from the usage perspective (usage-PMI), and then develop a Bayesian inference framework over the usage-flow graph to estimate the usage-PMI. Extensive experiments over three large CQA corpora show strong performance advantage over the baselines adopted from paraphrase identification task. We release 885,000 paraphrase templates of high quality discovered by our proposed PMI decomposition model, and the data is available in site https://github.com/Para-Questions/Intention\_template\_discovery.
Denghao Ma, Kevin Chen-Chuan Chang, Yueguo Chen, Xueqiang Lv
CIKM4
2023 CFNet: Head detection network based on multi-layer feature fusion and attention mechanism
abstract
Abstract Recently, head detection has been widely used in target detection, which has a great application value for improving security prevention and control in public places, as well as enhancing target tracking and identification in national defense, criminal investigation, and other fields. However, detecting small targets accurately at long distances is very difficult, and current methods often lack optimization of multi‐resolution features. Therefore, the authors propose a one‐stage detection network CFNet (cross‐layer feature fusion and fusion weight attention network), in which a fusion weight attention mechanism module (FWAM) is proposed to give different weights to the fused features in order to distinguish the importance of different features. The module increases the weights of features that contain strong information so that the fused features are focused on feature points that are beneficial for optimal head detection. Meanwhile, a cross‐layer feature fusion module is proposed to fuse information from different resolution feature maps to compensate for the decrease in detection accuracy caused by the omission of information features at low resolution, and a connection network for contextual information fusion is constructed, while weight parameter value settings are introduced to optimize the detection effect after fusion of different resolution features. In order to better reflect the effectiveness of the network, the experiments are performed on the SCUT‐HEAD PartA dataset and the Brainwash dataset; the results show that the network the authors proposed is better than the existing comparison methods, which proves the robustness and effectiveness of the network.
Xichang Wang, Xueqiang Lv
IET Image Process.4
2023 LiteDEKR: End-to-end lite 2D human pose estimation network
abstract
Abstract The 2D human pose estimation plays an important role in human‐computer interaction and action recognition. Although the method based on high‐resolution network has superior performance, there is still room for improvement in terms of speed and lightweight. Here, a LiteDEKR, a 2D pose estimation method that combines lightweight and accuracy, is proposed by designing a lightweight network based on DEKR and constructing two scientifically valid loss functions. The method, constructs a multi‐instance bias regression loss that matches the true distribution of keypoint bias, improves the accuracy of bias regression, and constructs a keypoint similarity loss with the object keypoint similarity index of keypoints as the optimization objective to achieve end‐to‐end training of the network. In addition, this paper has designed a lightweight DEKR, using LitePose as the backbone network. With the optimization of the above two loss functions, LiteDEKR not only achieves lightweight but also has high accuracy. Comparative experiments on the COCO and CrowdPose datasets show that compared to the current state‐of‐the‐art Contextual Instance Decoupling, LiteDEKR achieves a similar accuracy with only 10% of its network complexity. It also shows better robustness to low‐resolution input images.
Xueqiang Lv, Lianghai Tian, Zangtai Cai
IET Image Process.1
2023 HBert: A Long Text Processing Method Based on BERT and Hierarchical Attention Mechanisms
abstract
With the emergence of a large-scale pre-training model based on the transformer model, the effect of all-natural language processing tasks has been pushed to a new level. However, due to the high complexity of the transformer's self-attention mechanism, these models have poor processing ability for long text. Aiming at solving this problem, a long text processing method named HBert based on Bert and hierarchical attention neural network is proposed. Firstly, the long text is divided into multiple sentences whose vectors are obtained through the word encoder composed of Bert and the word attention layer. And the article vector is obtained through the sentence encoder that is composed of transformer and sentence attention. Then the article vector is used to complete the subsequent tasks. The experimental results show that the proposed HBert method achieves good results in text classification and QA tasks. The F1 value is 95.7% in longer text classification tasks and 75.2% in QA tasks, which are better than the state-of-the-art model longformer.
Xueqiang Lv, Zhaonan Liu, Xindong You
Int. J. Semantic Web Inf. Syst.1
2023 Research on the Generation of Patented Technology Points in New Energy Based on Deep Learning
abstract
Effective extraction of patent technology points in new energy fields is profitable, which motivates technological innovation and facilitates patent transformation and application. However, since patent data exists the ununiform distribution of technology points information, long length of term, and long sentences, technology point extraction faces the dilemmas of poor readability and logic confusion. To mitigate these problems, the article proposes a method to generate patent technology points called IGPTP—a two-stage strategy, which fuses the advantage of extractive and generative ways. IGPTP utilizes the RoBERTa+CNN model to obtain the key sentences of text and takes the output as input of UNILM (unified pre-trained language model). Simultaneously, it takes a multi-strategies integration technique to enhance the quality of patent technology points by combining the copy mechanism and external knowledge guidance model. Substantial experimental results manifest that IGPTP outperforms the current mainstream models, which can generate more coherent and richer text.
Haixiang Yang, Xindong You, Xueqiang Lv
Int. J. Semantic Web Inf. Syst.3
2022 MQDS: An energy saving scheduling strategy with diverse QoS constraints towards reconfigurable cloud storage systems
Xindong You, Dawei Sun 0001, Xueqiang Lv, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.3
2021 SS6: Online Short-Code RAID-6 Scaling by Optimizing New Disk Location and Data Migration
abstract
Abstract Thanks to excellent reliability, availability, flexibility and scalability, redundant arrays of independent (or inexpensive) disks (RAID) are widely deployed in large-scale data centers. RAID scaling effectively relieves the storage pressure of the data center and increases both the capacity and I/O parallelism of storage systems. To regain load balancing among all disks including old and new, some data usually are migrated from old disks to new disks. Owing to unique parity layouts of erasure codes, traditional scaling approaches may incur high migration overhead on RAID-6 scaling. This paper proposes an efficient approach based Short-Code for RAID-6 scaling. The approach exhibits three salient features: first, SS6 introduces $\tau $ to determine where new disks should be inserted. Second, SS6 minimizes migration overhead by delineating migration areas. Third, SS6 reduces the XOR calculation cost by optimizing parity update. The numerical results and experiment results demonstrate that (i) SS6 reduces the amount of data migration and improves the scaling performance compared with Round-Robin and Semi-RR under offline, (ii) SS6 decreases the total scaling time against Round-Robin and Semi-RR under two real-world I/O workloads (iii) the user average response time of SS6 is better than the other two approaches during scaling and after scaling.
Xindong You, Xueqiang Lv
Comput. J.3
2021 K-ear: Extracting data access periodic characteristics for energy-aware data clustering and storing in cloud storage systems
abstract
Abstract Rapid increase in energy consumption is a serious problem in cloud storage systems. Data accessed in large‐scale storage systems usually exhibit temporal and spatial characteristics, which make it possible to reduce energy consumption by clustering data with similar access characteristics for storage in the same zone of cloud storage systems. Existing works usually only focus on the frequency of data access. However, widely existing phenomena show data access with seasonal and tidal characteristics in cloud storage systems. The seasonal and tidal characteristics of data access are extracted thoroughly in this paper. According to the extracted data access characteristics, energy‐aware data clustering through a machine learning algorithm (K‐ear) is proposed. K‐ear classifies data into five seasonal categories according to their seasonal access characteristics and then classifies every seasonal category into three tidal categories according to its tidal access characteristics. The 15 classified categories are stored in different storage zones with different energy and performance modes. Simulation experiments using CloudSimDisk with the constructed mathematic models demonstrate that the proposed K‐ear algorithm is more energy‐efficient than the default data clustering algorithms in Hadoop and the classical data clustering storage strategy according to the data access frequency (Striping‐Based Energy‐Aware Strategy).
Xindong You, Dawei Sun 0001, Xunyun Liu, Xueqiang Lv, Rajkumar Buyya
Concurr. Comput. Pract. Exp.5
2021 HS6: An Efficient H-Code RAID-6 Scaling by Optimizing Data Migrating and Parity Updating
Xindong You, Xueqiang Lv, Muyuan Li
J. Supercomput.3
2020 Sentiment Analysis of Film Reviews Based on Deep Learning Model Collaborated with Content Credibility Filtering
Xindong You, Xueqiang Lv, Shangqian Zhang, Dawei Sun 0001, Shang Gao 0003
CollaborateCom (1)2
2020 Distant Supervised Relation Extraction via DiSAN-2CNN on a Feature Level
abstract
At present, the mainstream distant supervised relation extraction methods existed problems: the coarse granularity for coding the context feature information; the difficulty in capturing the long-term dependency in the sentence, and the difficulty in coding prior knowledge of structures are major issues. To address these problems, we propose a distant supervised relation extraction model via DiSAN-2CNN on feature level, in which multi-dimension self-attention mechanism is utilized to encode the features of the words and DiSAN-2CNN is used to encode the sentence to obtain the long-term dependency, the prior knowledge of the structure, the time sequence, and the entity dependence in the sentence. Experiments conducted on the NYT-Freebase benchmark dataset demonstrate that the proposed DiSAN-2CNN on a feature level model achieves better performance than the current two state-of-art distant supervised relation extraction models PCNN+ATT and ResCNN-9, and it has d generalization ability with the least artificial feature engineering.
Xueqiang Lv, Huixin Hou, Xindong You, Junmei Han
Int. J. Semantic Web Inf. Syst.1
2019 Relation Extraction Toward Patent Domain Based on Keyword Strategy and Attention+BiLSTM Model (Short Paper)
Xueqiang Lv, Xiangru Lv, Xindong You, Zhian Dong, Junmei Han
CollaborateCom1
2019 Momentum Based on Adaptive Bold Driver
abstract
The momentum-based stacked attention networks (SANs) is one of the best models for image question answering. However, we find that it is easy to fall into the local optimal solution, which results in the higher question answering error rate. To solve the problem, we propose adaptive bold driver (ABD). The experimental results and analysis show that it outperforms the state-of-the-art global learning rate adaptive algorithm in the local learning rate adaptive stochastic gradient descent (SGD). It is deeply integrated with momentum, and we propose momentum based on ABD (MABD). The experimental results show that its accuracy is 2.33% higher than the baseline (momentum), 2.54% higher than momentum based on bold driver, and 1.80% higher than the annealing-based momentum. The experimental analysis proves that it is the state-of-the-art optimization algorithm in the SANs-based image question answering and it has effectiveness, significance, generalization performance, and promotional value.
Shengdong Li, Xueqiang Lv
ICME2
2016 Patent Subject Words Extraction Based on Integrated Strategy Method
abstract
The extraction of patent subject words is the principal task in the patent analysis. In this paper, we apply the technology of information mining to the patent literature, and propose a new method for the automatic extraction of patent subject words. According to the characteristics of patent theme, a method which combined the reverse combinational words and Mutual Information is adopted to extract the candidate words, and then the candidate words are filtered, finally according to the number of lexical units, the frequency and the position in the title, it puts forward a computing method of the theme degree and treats the biggest value as the subject word. The extraction method is tested on the patent data set in the domain of new energy automobile, and the precision rate is 85.65%, the recall rate is 82.4%. Experimental results show that it has a better effect, and has positive significance for the patent mining, the analysis and prediction of the patent technology.
Liya Zhu, Xueqiang Lv, Liping Xu
ISPDC2
2016 Integrating multiple types of features for event identification in social images
Xiaoming Zhang 0001, Zhoujun Li 0001, Xueqiang Lv, Xiaoming Chen 0007
Multim. Tools Appl.3
2016 Geographical Topics Learning of Geo-Tagged Social Images
abstract
With the availability of cheap location sensors, geotagging of images in online social media is very popular. With a large amount of geo-tagged social images, it is interesting to study how these images are shared across geographical regions and how the geographical language characteristics and vision patterns are distributed across different regions. Unlike textual document, geo-tagged social image contains multiple types of content, i.e., textual description, visual content, and geographical information. Existing approaches usually mine geographical characteristics using a subset of multiple types of image contents or combining those contents linearly, which ignore correlations between different types of contents, and their geographical distributions. Therefore, in this paper, we propose a novel method to discover geographical characteristics of geo-tagged social images using a geographical topic model called geographical topic model of social images (GTMSIs). GTMSI integrates multiple types of social image contents as well as the geographical distributions, in which image topics are modeled based on both vocabulary and visual features. In GTMSI, each region of the image would have its own topic distribution, and hence have its own language model and vision pattern. Experimental results show that our GTMSI could identify interesting topics and vision patterns, as well as provide location prediction and image tagging.
Xiaoming Zhang 0001, Shufan Ji, Senzhang Wang, Zhoujun Li 0001, Xueqiang Lv
IEEE Trans. Cybern.5
2015 Research on Semantic Disambiguation in Treebank
Xueqiang Lv, Yunfang Wu
APWeb2
2015 Location Prediction of Social Images via Generative Model
abstract
The vast amount of geo-tagged social images has attracted great attention in research of predicting location using the plentiful content of images, such as visual content and textual description. Most of the existing researches use the text-based or vision-based method to predict location. There still exists a problem: how to effectively exploit the correlation between different types of content as well as their geographical distributions for location prediction. In this paper, we propose to predict image location by learning the latent relation between geographical location and multiple types of image content. In particularly, we propose a geographical topic model GTMSI (geographical topic model of social image) to integrate multiple types of image content as well as the geographical distributions. In GTMI, image topic is modeled on both text vocabulary and visual feature. Each region has its own distribution over topics and hence has its own language model and vision pattern. The location of a new image is estimated based on the joint probability of image content and similarity measure on topic distribution between images. Experiment results demonstrate the performance of location prediction based on GTMSI.
Xiaoming Zhang 0001, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xueqiang Lv
ICMR5
2015 Multi-sentence Question Segmentation and Compression for Question Answering
abstract
We present a multi-sentence question segmentation strategy for community question answering services to alleviate the complexity of long sentences. We develop a complete scheme and make a solution to complex-question segmentation, including a question detector to extract question sentences, a question compression process to remove duplicate information, and a graph model to segment multi-sentence questions. In the graph model, we train a SVM classifier to compute the initial weight and we calculate the authority of a vertex to guide the propagating. The experimental results show that our method gets a good balance between completeness and redundancy of information, and significantly outperforms state-of-the-art methods.
Yixiu Wang, Yunfang Wu, Xueqiang Lv
NLPCC3
2015 Improving Chinese Dependency Parsing with Lexical Semantic Features
abstract
Lexical semantic information plays an important role in supervised dependency parsing. In this paper, we add lexical semantic features to the feature set of a parser, obtaining improvements on the Penn Chinese Treebank. We extract semantic categories of words from HowNet, and use them as semantic information of words. Moreover, we investigate the method to compute semantic similarity between Chinese compound words, and obtain semantic information of words which did not record in HowNet. Our experiments show that unlabeled attachment scores can increase by 1.29%.
Lvexing Zheng, Houfeng Wang, Xueqiang Lv
NLPCC3
2014 Automatic Recognition of Chinese Location Entity
Xueqiang Lv, Kehui Liu
NLPCC2
2013 i, Poet: Automatic Chinese Poetry Composition through a Generative Summarization Framework under Constrained Optimization
Rui Yan 0001, Mirella Lapata, Shou-De Lin, Xueqiang Lv, Xiaoming Li 0001
IJCAI5
2013 Semantic v.s. Positions: Utilizing Balanced Proximity in Language Model Smoothing for Information Retrieval
Rui Yan 0001, Mirella Lapata, Shou-De Lin, Xueqiang Lv, Xiaoming Li 0001
IJCNLP5
2010 Research and Application to Automatic Indexing
Shuicai Shi, Xueqiang Lv, Yuqin Li
ISNN (2)3
2007 Design and Realization of Advertisement Promotion Based on the Content of Webpage
Shuicai Shi, Xueqiang Lv
KSEM3
2006 A Comparative Study on Representing Units in Chinese Text Clustering
Shiwen Yu, Xueqiang Lv, Shuicai Shi, Shibin Xiao
KSEM3