VLDB 2026 Research / reviewers in the wild / expert
Qun Li 0002
dblp:42/6066-2
· DBLP profile ↗
36ranked-venue papers
19as first author
20since 2021 · last 2026
0000-0002-8034-6030ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 6 since 2021Computer networks · 8 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TSMnet: Two-step separation pipeline based on threshold shrinkage memory network for weakly-supervised video anomaly detection
Qun Li 0002, Xinping Gao, Bir Bhanu |
Pattern Recognit. Lett. | 1 |
| 2026 | PE-ViT: Parameter-efficient vision transformer with dimension-adaptive experts and economical attention
Qun Li 0002, Jiru He, Tiancheng Guo, Xinping Gao, Bir Bhanu |
Pattern Recognit. Lett. | 1 |
| 2026 | Toward Generative Understanding: Incremental Few-Shot Semantic Segmentation With Diffusion ModelsabstractIncremental Few-shot Semantic Segmentation (iFSS) aims to learn novel classes with limited samples while preserving segmentation capability for base classes, addressing the challenge of continual learning of novel classes and catastrophic forgetting of previously seen classes. Existing methods mainly rely on techniques such as knowledge distillation and background learning, which, while partially effective, still suffer from issues such as feature drift and limited generalization to real-world novel classes, primarily due to a bidirectional coupling bottleneck between the learning of base classes and novel classes. To address these challenges, we propose, for the first time, a diffusion-based generative framework for iFSS. Specifically, we bridge the gap between generative and discriminative tasks through an innovative binary-to-RGB mask mapping mechanism, enabling pre-trained diffusion models to focus on target regions via class-specific semantic embedding optimization while sharpening foreground-background contrast with color embeddings. A lightweight post-processor then refines the generated images into high-quality binary masks. Crucially, by leveraging diffusion priors, our framework avoids complex training strategies. The optimization of class-specific semantic embeddings decouples the embedding spaces of base and novel classes, inherently preventing feature drift, mitigating catastrophic forgetting, and enabling rapid novel-class adaptation. Experimental results show that our method achieves state-of-the-art performance on the PASCAL- $5^{i}$ and COCO- $20^{i}$ datasets using much less data than other methods, and exhibiting competitive results in cross-domain few-shot segmentation tasks. Project page: https://ifss-diff.github.io/. Qun Li 0002, Fu Xiao 0001, Na Zhao 0004, Bir Bhanu |
IEEE Trans. Image Process. | 1 |
| 2025 | Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View ClusteringabstractMulti-view clustering aims to enhance clustering performance by leveraging information from diverse sources. However, its practical application is often hindered by a barrier: the lack of correspondences across views. This paper focuses on the understudied problem of fully incomplete multi-view clustering (FIMC), a scenario where existing methods fail due to their reliance on partial alignment. To address this problem, we introduce the Contrastive Prototype Matching Network (CPMN), a novel framework that establishes a new paradigm for cross-view alignment based on matching high-level categorical structures. Instead of aligning individual instances, CPMN performs a more robust cluster prototype alignment. CPMN first employs a correspondence-free graph contrastive learning approach, leveraging mutual $k$-nearest neighbors (MNN) to uncover intrinsic data structures and establish initial prototypes from entirely unpaired views. Building on the prototypes, we introduce a cross-view prototype graph matching stage to resolve category misalignment and forge a unified clustering structure. Finally, guided by this alignment, we devise a prototype-aware contrastive learning mechanism to promote semantic consistency, replacing the reliance on the initial MNN-based structural similarity. Extensive experiments on benchmark datasets demonstrate that our method significantly outperforms various baselines and ablation variants, validating its effectiveness. Yiming Wang 0007, Qun Li 0002, Dongxia Chang, Jie Wen 0001, Hua Dai 0003, Fu Xiao 0001, Yao Zhao 0001 |
NeurIPS | 2 |
| 2025 | PS-CoT-Adapter: adapting plan-and-solve chain-of-thought for ScienceQA
Qun Li 0002, Fu Xiao 0001, Yiming Wang 0007, Xinping Gao, Bir Bhanu |
Sci. China Inf. Sci. | 1 |
| 2025 | Spatial-temporal multi-scale interaction for few-shot video summarization
Qun Li 0002, Zhuxi Zhan, Yanchao Li 0001, Bir Bhanu |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A Category-Driven Contrastive Recovery Network for Double Incomplete Multi-View Multi-Label ClassificationabstractIn the field of multi-view multi-label learning, the challenges of incomplete views and missing labels are prevalent due to the complexity of manual labeling and data acquisition errors. These challenges significantly reduce the quality of latent representations and hinder prediction by multi-label classification. To address this issue, we propose a novel Category-driven Semi-supervised Contrastive Recovery (CSCR) framework in this study. Our framework aims to fully integrate existing label information into incomplete representation learning and classification. Specifically, to address the limitations posed by incomplete views and labels, we construct a label coincidence matrix based on existing labels, which serves as a similarity matrix in subsequent semi-supervised contrastive learning and multi-view classification. By leveraging this matrix, we design a semi-supervised multi-view contrastive learning module, which constructs sample pairs on the basis of inter-view correspondences and label similarity. It learns discriminative latent representations without the need for data augmentation. A weighted multi-label classification module is subsequently employed to integrate the predictions from each view to obtain the final classification result. Experimental evaluations on five challenging datasets demonstrate the superiority of our model over existing state-of-the-art methods. Yiming Wang 0007, Qun Li 0002, Dongxia Chang, Jie Wen 0001, Fu Xiao 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Dynamic context modeling based lightweight high-resolution network for dense prediction
Baoquan Sun, Qun Li 0002, Ziyi Zhang 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Co-GZSL: Feature Contrastive Optimization for Generalized Zero-Shot LearningabstractAbstract Generalized Zero-Shot Learning (GZSL) learns from only labeled seen classes during training but discriminates both seen and unseen classes during testing. In GZSL tasks, most of the existing methods commonly utilize visual and semantic features for training. Due to the lack of visual features for unseen classes, recent works generate real-like visual features by using semantic features. However, the synthesized features in the original feature space lack discriminative information. It is important that the synthesized visual features should be similar to the ones in the same class, but different from the other classes. One way to solve this problem is to introduce the embedding space after generating visual features. Following this situation, the embedded features from the embedding space can be inconsistent with the original semantic features. For another way, some recent methods constrain the representation by reconstructing the semantic features using the original visual features and the synthesized visual features. In this paper, we propose a hybrid GZSL model, named feature Contrastive optimization for GZSL (Co-GZSL), to reconstruct the semantic features from the embedded features, which ensures that the embedded features are close to the original semantic features indirectly by comparing reconstructed semantic features with original semantic features. In addition, to settle the problem that the synthesized features lack discrimination and semantic consistency, we introduce a Feature Contrastive Optimization Module (FCOM) and jointly utilize contrastive and semantic cycle-consistency losses in the FCOM to strengthen the intra-class compactness and the inter-class separability and to encourage the model to generate semantically consistent and discriminative visual features. By combining the generative module, the embedding module, and the FCOM, we achieve Co-GZSL. We evaluate the proposed Co-GZSL model on four benchmarks, and the experimental results indicate that our model is superior over current methods. Code is available at: https://github.com/zhanzhuxi/Co-GZSL . Qun Li 0002, Zhuxi Zhan, Yaying Shen, Bir Bhanu |
Neural Process. Lett. | 1 |
| 2023 | ESSL: Enhanced Spatio-Temporal Self-Selective Learning Framework for Unsupervised Video Anomaly DetectionabstractUnsupervised Video Anomaly Detection (UVAD) utilizes completely unlabeled videos for training without any human intervention. Due to the existence of unlabeled abnormal videos in the training data, the performance of UVAD has a large gap compared with semi-supervised VAD, which only uses normal videos for training. To address the problem of insufficient ability of the existing UVAD methods to learn normality and reduce the negative impact of abnormal events, this paper proposes a novel Enhanced Spatio-temporal Self-selective Learning (ESSL) framework for UVAD. This framework is designed for capturing both the appearance and motion features through effective network structures by solving the spatial and temporal jigsaw puzzles. Specially, we develop a Self-selective Learning Module (SLM) for UVAD, which prevents the model learning abnormal features and enhances the model by selecting normal features. Experimental results on three benchmark datasets show that the proposed method not only surpasses the state-of-the-art UVAD works, but also achieves the performance comparable to the classic semi-supervised methods for video anomaly detection that needs normal videos selected manually. Code is available at: https://github.com/xusuger/ESSL. Qun Li 0002, Xubei Pan, Fu Xiao 0001, Bir Bhanu |
ECAI | 1 |
| 2023 | Lite-FENet: Lightweight multi-scale feature enrichment network for few-shot segmentation
Qun Li 0002, Baoquan Sun, Bir Bhanu |
Knowl. Based Syst. | 1 |
| 2023 | HRNeXt: High-Resolution Context Network for Crowd Pose EstimationabstractOcclusion handling in crowded scenes is an intractable challenge for human pose estimation. To address this problem, we propose two novel feed-forward network structures named Global Feed-Forward Network (GFFN) and Dynamic Feed-Forward Network (DFFN), which are specifically designed for image-based tasks to capture both local and global contextual information within intermediate features and update feature representations with high adaptability for occlusions. By exploiting the context modeling ability of the proposed GFFN and DFFN, we present a novel backbone network, namely High-Resolution Context Network (HRNeXt), which learns high-resolution representations with abundant contextual information to better estimate poses of occluded human bodies. Compared to state-of-the-art pose estimation networks, our HRNeXt absorbs advantages of convolution operation and attention mechanism, and it is more efficient in terms of training data sizes, network parameters and computational costs. Experimental results show that our HRNeXt significantly outperforms state-of-the-art backbone networks on challenging pose estimation datasets with high occurrence of crowds and occlusions. Qun Li 0002, Ziyi Zhang 0001, Fu Xiao 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | FedDyn: A dynamic and efficient federated distillation approach on Recommender SystemabstractFederated Learning (FL) is a popular distributed machine learning paradigm that enables devices to work together to train a centralized model without transmitting raw data. However, when the model becomes complex, mobile devices’ communication overhead can be unacceptably large in traditional FL methods. To address this problem, Federated Distillation (FD) is proposed as a federated version of knowledge distillation. Most of the recent FD methods calculate the model output (logits) of each client as the local knowledge on a public proxy dataset and do distillation with the average of the clients’ logits on the server side. Nevertheless, these FD methods are not robust and perform poorly in the non-IID (data is nonindependent and non-identically distributed) scenario such as Federated Recommendation (FR). In order to eliminate the non-IID problem and apply FD in FR, we proposed a novel method named FedDyn to construct a proxy dataset and extract local knowledge dynamically in this paper. In this method, we replaced the average strategy with focus distillation to strengthen reliable knowledge, which solved the non-IID problem that the local model has biased knowledge. The average strategy is a dilution and perturbation of knowledge since it treats reliable and unreliable knowledge equally important. In addition, to prevent inference of private user information from local knowledge, we used a method like local differential privacy techniques to protect this knowledge on the client side. The experimental results showed that our method has a faster convergence speed and lower communication overhead than the baselines on three datasets, including MovieLens-10OK, MovieLens-IM and Pinterest. Xuandong Chen, Qun Li 0002 |
ICPADS | 4 |
| 2022 | Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose EstimationabstractA high-resolution network exhibits remarkable capability in extracting multi-scale features for human pose estimation, but fails to capture long-range interactions between joints and has high computational complexity. To address these problems, we present a Dynamic lightweight High-Resolution Network (Dite-HRNet), which can efficiently extract multi-scale contextual information and model long-range spatial dependency for human pose estimation. Specifically, we propose two methods, dynamic split convolution and adaptive context modeling, and embed them into two novel lightweight blocks, which are named dynamic multi-scale context block and dynamic global context block. These two blocks, as the basic component units of our Dite-HRNet, are specially designed for the high-resolution networks to make full use of the parallel multi-resolution architecture. Experimental results show that the proposed network achieves superior performance on both COCO and MPII human pose estimation datasets, surpassing the state-of-the-art lightweight networks. Code is available at: https://github.com/ZiyiZhang27/Dite-HRNet. Qun Li 0002, Ziyi Zhang 0001, Fu Xiao 0001, Bir Bhanu |
IJCAI | 1 |
| 2022 | Anomaly Detection in Surveillance Videos via Memory-augmented Frame PredictionabstractAnomaly detection in surveillance videos is a challenging task in computer vision, and can be defined as the detection of actions or events that do not conform to the expected behaviors. Most of the existing methods solve the task by minimizing the reconstruction errors between the ground-truth video frames and their reconstructed frames. However, these methods sometimes reconstruct anomalies well that results in high false detections and a decrease of the performance. Therefore, we propose a frame prediction method which is based on a memory-augmented scheme for anomaly detection. Our method regards anomaly detection as a frame prediction task, and uses a generative network to achieve the frame prediction. For generating high quality video frames, we embed a memory module into the generative network, which effectively improves the feature representation of normal events and reduces the representation of abnormal events. In addition, we adapt an attention mechanism to model the interdependence between feature channels. In order to evaluate our method, we introduce a new anomaly detection dataset that consists of real and multi-scene surveillance videos. Extensive experiments on our dataset and publicly available datasets validate the effectiveness and robustness of our proposed method. Qun Li 0002, Yaying Shen, Ziyi Zhang 0001 |
IJCNN | 2 |
| 2022 | Attention-based anomaly detection in multi-view surveillance videos
Qun Li 0002, Fu Xiao 0001, Bir Bhanu |
Knowl. Based Syst. | 1 |
| 2022 | Inner Knowledge-based Img2Doc Scheme for Visual Question AnsweringabstractVisual Question Answering (VQA) is a research topic of significant interest at the intersection of computer vision and natural language understanding. Recent research indicates that attributes and knowledge can effectively improve performance for both image captioning and VQA. In this article, an inner knowledge-based Img2Doc algorithm for VQA is presented. The inner knowledge is characterized as the inner attribute relationship in visual images. In addition to using an attribute network for inner knowledge-based image representation, VQA scheme is associated with a question-guided Doc2Vec method for question–answering. The attribute network generates inner knowledge-based features for visual images, while a novel question-guided Doc2Vec method aims at converting natural language text to vector features. After the vector features are extracted, they are combined with visual image features into a classifier to provide an answer. Based on our model, the VQA problem is resolved by textual question answering. The experimental results demonstrate that the proposed method achieves superior performance on multiple benchmark datasets. Qun Li 0002, Fu Xiao 0001, Bir Bhanu, Biyun Sheng, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Legitimate Eavesdropping with Wireless Powered Proactive Full-duplex EavesdroppersabstractIn this paper, legitimate eavesdropping in a point-to-point suspicious communication network with multiple wireless powered full-duplex legitimate eavesdroppers is investigated. The eavesdroppers are assumed to adopt the power splitting technique to coordinate energy harvesting and information eavesdropping, and use the harvested energy to proactively jam the suspicious communication. For both collusive and non-collusive eavesdroppers, the successful eavesdropping probability is maximized by optimizing the power splitting ratio at each eavesdropper. Heuristic algorithms are proposed to solve the optimization problems. Simulation results are provided to confirm the effectiveness of the proposed algorithms. It is shown that the proposed algorithms outperform the reference algorithms without proactive jamming, especially for non-collusive eavesdroppers with a high energy harvesting efficiency, a high transmit power of the suspicious user, or a large number of eavesdroppers. Qun Li 0002 |
WCNC | 1 |
| 2021 | Joint User Selection and Power Control for Secure Communication in Multicast NetworksabstractIn this paper, we investigate physical layer security of a multicast network consisting of multiple legitimate users in the presence of an eavesdropper. Besides information receivers, the legitimate users are assumed to be able to act as jammers for interfering with the eavesdropper to improve secrecy. The joint user selection and power control problem for maximizing the average sum secrecy rate of the multicast network under the transmit power constraints is investigated with perfect channel state information (CSI) or with unknown CSI related to the eavesdropper, and heuristic algorithms are proposed correspondingly. Simulation results are presented to verify the proposed algorithms. It is shown that the proposed algorithms significantly outperform the reference algorithms in terms of average sum secrecy rate. Qun Li 0002 |
WCNC | 1 |
| 2021 | Cooperative Resource Allocation for Computation Offloading in Mobile-Edge Computing NetworksabstractThis paper considers a mobile-edge computing (MEC) network with multiple cooperative MEC servers. It is assumed that the MEC server can transfer a fraction of user task to other MEC servers if its computation capacity is insufficient. The problem of cooperative resource allocation for computation offloading to minimize the energy consumption of all users under the task latency constraint, user transmission rate constraint and the MEC computation capacity constraint is investigated, and a suboptimal iterative algorithm is proposed based on alternating optimization and convex optimization. It is shown that the proposed algorithm greatly outperforms the algorithm without cooperation and the benchmark algorithm. Qun Li 0002, Hanqin Shao |
WCNC | 1 |
| 2020 | Discriminative Multi-View Subspace Feature Learning for Action RecognitionabstractAlthough deep features have achieved the state-of-the-art performance in action recognition recently, the hand-crafted shallow features still play a critical role in characterizing human actions for taking advantage of visual contents in an intuitive way such as edge features. Therefore, the shallow features can serve as auxiliary visual cues supplementary to deep representations. In this paper, we propose a discriminative subspace learning model (DSLM) to explore the complementary properties between the hand-crafted shallow feature representations and the deep features. As for the RGB action recognition, this is the first work attempting to mine multi-level feature complementaries by the multi-view subspace learning scheme. To sufficiently capture the complementary information among heterogeneous features, we construct the DSLM by integrating the multi-view reconstruction error and classification error into an unified objective function. To be specific, we first use Fisher Vector to encode improved dense trajectories (iDT+FV) for shallow representations and two-stream convolutional neural network models (T-CNN) for generating deep features. Moreover, the presented DSLM algorithm projects multi-level features onto a shared discriminative subspace with the complementary information and discriminating capacity simultaneously incorporated. Finally, the action types of test samples are identified by the margins from the learned compact representations to the decision boundary. The experimental results on three datasets demonstrate the effectiveness of the proposed method. Biyun Sheng, Jun Li 0033, Fu Xiao 0001, Qun Li 0002, Wankou Yang, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Resource allocation in cognitive wireless powered communication networks with wirelessly powered secondary users and primary users
Ding Xu 0001, Qun Li 0002 |
Sci. China Inf. Sci. | 2 |
| 2019 | Cooperative resource allocation in cognitive wireless powered communication networks with energy accumulation and deadline requirements
Ding Xu 0001, Qun Li 0002 |
Sci. China Inf. Sci. | 2 |
| 2019 | Semantic Concept Network and Deep Walk-based Visual Question AnsweringabstractVisual Question Answering (VQA) is a hot-spot in the intersection of computer vision and natural language processing research and its progress has enabled many in high-level applications. This work aims to describe a novel VQA model based on semantic concept network construction and deep walk. Extracting visual image semantic representation is a significant and effective method for spanning the semantic gap. Moreover, current research has shown that co-occurrence patterns of concepts can enhance semantic representation. This work is motivated by the challenge that semantic concepts have complex interrelations and the relationships are similar to a network. Therefore, we construct a semantic concept network adopted by leveraging Word Activation Forces (WAFs), and mine the co-occurrence patterns of semantic concepts using deep walk. Then the model performs polynomial logistic regression on the basis of the extracted deep walk vector along with the visual image feature and question feature. The proposed model effectively integrates visual and semantic features of the image and natural language question. The experimental results show that our algorithm outperforms competitive baselines on three benchmark image QA datasets. Furthermore, through experiments in image annotation refinement and semantic analysis on pre-labeled LabelMe dataset, we test and verify the effectiveness of our constructed concept network for mining concept co-occurrence patterns, sensible concept clusters, and hierarchies. Qun Li 0002, Fu Xiao 0001, Xianzhong Long, Xiaochuan Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Recurrent neural system with minimum complexity: A deep learning perspective
Xiaochuan Sun, Tao Li 0001, Yingqi Li, Qun Li 0002 |
Neurocomputing | 4 |
| 2017 | Price-based time and energy allocation in cognitive radio multiple access networks with energy harvesting
Ding Xu 0001, Qun Li 0002 |
Sci. China Inf. Sci. | 2 |
| 2017 | Improving physical-layer security for primary users in cognitive radio networksabstractIn this study, the authors investigate the physical‐layer security in a cognitive radio network where both the secondary user (SU) and the primary user (PU) are facing security threats from the malicious eavesdroppers. To protect the PU, the SU acts as a friendly jammer to interfere with the eavesdroppers by splitting a certain portion of the transmit power for sending the jamming noise. The problem of optimising SU scheduling, power allocation and power splitting ratio to maximise the SU ergodic secrecy rate subject to the PU secrecy outage constraint with imperfect channel state information available at the SU is investigated based on the dual optimisation method. In addition, a greedy algorithm is also proposed for minimising the PU secrecy outage probability. Simulation results indicate that the proposed algorithms are effective in improving the PU secrecy performance in terms of secrecy outage probability as well as providing secure communications for the SU. Ding Xu 0001, Qun Li 0002 |
IET Commun. | 2 |
| 2017 | Deep belief echo-state network and its application to time series prediction
Xiaochuan Sun, Tao Li 0001, Qun Li 0002, Yingqi Li |
Knowl. Based Syst. | 3 |
| 2015 | Optimal power allocation for cognitive radio networks with primary user secrecy rate loss constraintabstractThis paper considers a cognitive radio (CR) network in the presence of a malicious eavesdropper who attempts to receive confidential messages from a pair of primary users (PUs). Under the PU secrecy rate loss constraint and the SU maximum transmit power constraint, a pair of secondary users (SUs) is proposed to interfere with the eavesdropper to improve the PU secrecy level and thus gain its own transmission opportunities. Then, the closed-form optimal power allocation strategy for the SU to maximize its transmission rate subject to the aforementioned constraints is derived. Extensions of the results to the scenarios with multiple eavesdroppers and multiple SUs are also presented, respectively. Numerous simulation results are illustrated to investigate the impacts of various system parameters on the SU transmission rate and the PU secrecy rate. Our results indicate that the PU secrecy rate improves significantly with the help of the SU transmission. Ding Xu 0001, Qun Li 0002 |
ICC | 2 |
| 2015 | Energy efficient joint chunk and power allocation for chunk-based multi-carrier cognitive radio networksabstractThis paper investigates the problem of energy efficient resource allocation in a chunk-based multi-carrier cognitive radio (CR) network. Chunk-based resource allocation is adopted where subcarriers are grouped into chunks to be allocated to the secondary users (SUs). The objective is to maximize the energy efficiency of the CR network while also satisfying the transmit power constraint as well as the interference power constraint for protecting the primary user (PU). For this, based on Dinkelbach method and dual optimization method, an efficient iterative algorithm is proposed. The impacts of the interference power constraint, the transmit power constraint, number of subcarriers within the chunk and the channel coherence bandwidth on the performance of the proposed algorithm are examined by simulations. It is shown that the proposed algorithm not only converges fast but also achieves almost the same performance as the exhaustive search algorithm does. It is also shown that the performance of the proposed algorithm significantly improves compared to the max-sum-rate algorithm and the equal power allocation algorithm especially for large transmit power and interference power limits. In addition, it is shown that the proposed algorithm achieves higher energy efficiency with less number of subcarriers within the chunk. Ding Xu 0001, Qun Li 0002 |
WCNC | 2 |
| 2015 | Effective capacity region and power allocation for two-way spectrum sharing cognitive radio networks
Ding Xu 0001, Qun Li 0002 |
Sci. China Inf. Sci. | 2 |
| 2015 | Construction of semantic bootstrapping models for relation extraction
Chunyun Zhang, Weiran Xu, Zhanyu Ma, Sheng Gao 0001, Qun Li 0002, Jun Guo 0002 |
Knowl. Based Syst. | 5 |
| 2013 | Ordered histogram of shapemes: An ordered bag-of-features based shape descriptor for efficient shape matchingabstractIn this paper, we enhance the Shape Context-based descriptor, shapemes, by introducing an ordered bag-of-features model and dynamic programming. The proposed descriptor consists of a series of sub-histograms of shapemes, each of which represents a subset of sampled points. The division of the sampled points is based on their sequential positions on the contour of the shape, so the representation has intrinsic order and is therefore named ordered histogram of shapemes. Then dynamic programming is utilized for descriptor matching. The framework is effective and efficient owing to the following properties: 1) points division approach together with dynamic programming for invariance under the change of starting point, 2) Earth Mover's Distance for discriminative power, and 3) pre-caculated shapemes dissimilarity matrix for fast descriptor distance calculation. Experiments on standard shape database and real world application scenario demonstrate the effectiveness and efficiency of the descriptor and the matching framework. We make our code and experimental data publicly available for future reference. Lunshao Chai, Zhen Qin 0001, Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002 |
ICIP | 3 |
| 2013 | Representative reference-set and betweenness centrality for scene image categorizationabstractReference-based image classification approach introduces a reference-set for both image representation and dictionary learning. It significantly reduces the dimensionality of represented images and shows outstanding performance even with randomly selected reference images and simple distance measure. In this paper, we improve upon existing work with two major contributions. First, we show that a more representative reference-set contributes to better classification accuracy. To this end, we carefully adapt the K-means clustering algorithm in the feature space to select a distinguished reference-set. Second, in the image classification process, we propose to represent each image by measuring its betweenness centrality in a social network composed of the representative reference-set in each class, leading to a more coherent distance measure that considers the overall connectivity between the probe image and the reference-set. Extensive experiment results demonstrate that our proposed scheme achieves better performance than existing methods. Qun Li 0002, Zhen Qin 0001, Lunshao Chai, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
ICIP | 1 |
| 2013 | Reference-Based Scheme Combined With K-SVD for Scene Image CategorizationabstractA reference-based algorithm for scene image categorization is presented in this letter. In addition to using a reference-set for images representation, we also associate the reference-set with training data in sparse codes during the dictionary learning process. The reference-set is combined with the reconstruction error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. After dictionaries are constructed, Locality-constrained Linear Coding (LLC) features of images are extracted. Then, we represent each image feature vector using the similarities between the image and the reference-set, leading to a significant reduction of the dimensionality in the feature space. Experimental results demonstrate that our method achieves outstanding performance. Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
IEEE Signal Process. Lett. | 1 |
| 2012 | Codebook optimization using word activation forces for scene categorizationabstractVisual codebook based quantization of robust appearance descriptors extracted from local image patches is an effective means of capturing image statistics for texture analysis and natural scene classification. In this paper, based on the newly proposed statistics of word activation forces (WAFs), we optimize the codebook. Currently, codebooks are typically created from a set of training images using a clustering algorithm. However, these codebooks are often functionally limited due to redundancy. We show that WAFs can remove the redundancy efficiently. In the experiment, the proposed method achieved the state-of-the-art performance on the Caltech-101, fifteen natural scene categories and VOC2007 databases. The optimization method also offers insights into the success of several recently proposed images classification approaches, including vector quantization (VQ) coding in the Spatial Pyramid Matching (SPM), sparse coding SPM (ScSPM), and Locality-constrained Linear Coding (LLC). Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
ICIP | 1 |