VLDB 2026 Research / reviewers in the wild / expert
Qi Wang 0079
dblp:19/1924-79
· DBLP profile ↗
24ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-6269-0196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WheatGOAT: Generalizable object-aware tracker via discriminative region semantic learning for wheat ear counting
Xingcai Wu, Yaoxi Li, Ziang Zou, Ya Yu, G. M. A. D. Sirishantha, A. S. A. Salgadoeb, Gefei Hao, Qi Wang 0079 |
Adv. Eng. Informatics | 9 |
| 2026 | Quantifying expressive power in knowledge graph embeddings: An entropy-based metric framework
Panfeng Chen, Hui Li 0046, Qi Wang 0079, Xibin Wang |
Expert Syst. Appl. | 3 |
| 2026 | DHS-ViG: Dynamic hierarchical selective graph for comprehensive and robust feature perception
Chaojie Chen, Xingcai Wu, Yuanyuan Xiao, Peijia Yu, Qi Wang 0079 |
Neurocomputing | 5 |
| 2026 | AdaptiveMamba: Comprehensive visual representation learning with adaptive semantic perception
Chaojie Chen, Xingcai Wu, Wu Liu 0005, Qi Wang 0079 |
Knowl. Based Syst. | 4 |
| 2025 | A generalized CNN decision boundary theory based on non-Euclidean space
Yongjun Zhou, Qinglei Li, Yuanyuan Xiao, Qi Wang 0079 |
Appl. Intell. | 6 |
| 2025 | Relation Semantic Guidance and Entity Position Location for Relation ExtractionabstractAbstract Relation extraction is a research hot-spot in the field of natural language processing, and aims at structured knowledge acquirement. However, existing methods still grapple with the issue of entity overlapping, where they treat relation types as inconsequential labels, overlooking the fact that relation type has a great influence on entity type hindering the performance of these models from further improving. Furthermore, current models are inadequate in handling the fine-grained aspect of entity positioning, which leads to ambiguity in entity boundary localization and uncertainty in relation inference, directly. In response to this challenge, a relation extraction model is proposed, which is guided by relational semantic cues and focused on entity boundary localization. The model uses an attention mechanism to align relation semantics with sentence information, so as to obtain the most relevant semantic expression to the target relation instance. It then incorporates an entity locator to harness additional positional features, thereby, enhancing the capability of the model to pinpoint entity start and end tags. Consequently, this approach effectively alleviates the problem of entity overlapping. Extensive experiments are conducted on the widely used datasets NYT and WebNLG. The experimental results show that the proposed model outperforms the baseline ones in F1 scores of the two datasets, and the improvement margin is up to 5.50% and 2.80%, respectively. Panfeng Chen, Hui Li 0046, Xibin Wang, Aihua Yu, Xingzhi Deng, Qi Wang 0079 |
Data Sci. Eng. | 8 |
| 2025 | EMGE: Entities and Mentions Gradual Enhancement with semantics and connection modelling for document-level relation extraction
Panfeng Chen, Qi Wang 0079, Hui Li 0046, Xibin Wang, Aihua Yu, Xingzhi Deng |
Knowl. Based Syst. | 3 |
| 2025 | DRC: Discrete Representation Classifier With Salient Features via Fixed-PrototypeabstractImage classification models including convolutional neural networks (CNN) and vision transformers (ViT) commonly employ a fully connected (FC) layer as the classifier. However, the fully connected nature of FC brings large amounts of weight parameters, limits the efficiency of inference, tends to over-fit the training data, and struggles to learn distinct class weights. To solve these problems, we propose a discrete representation classifier (DRC), a generic parameter-free classifier that offers efficiency, robustness, and more discriminative categorization. Specifically, the DRC discards numerous unimportant features and focuses solely on the salient features which are reinforced during training and presented in short discrete form during inference. Unlike the way of learning pseudo-prototypes (weights) from data laden with complex patterns and noises in FC, the DRC introducing discriminative fixed-prototypes which are almost uniformly distributed across the high-dimensional feature space, thus helps the model to learn more distinct boundaries between categories. Further leveraging the advantage of DRC’s focus on salient features, we propose Salient-CAM, which is able to locate the most important region in image without the need for weighting feature maps. The experiments demonstrate that simply replacing the model’s classifier from FC to DRC can lead to a significant acceleration in the whole model’s inference and a more robust classification. Additionally, the proposed Salient-CAM exhibits excellent object localization ability in complex natural scenes. Qinglei Li, Qi Wang 0079, Yongbin Qin, Xingcai Wu, Shiming Chen 0002, Wu Liu 0005, Yong-Jin Liu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation ModelabstractDepth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains challenging in robust generalization due to limited data resources and spectral differences between thermal and RGB images. In this paper, we present a self-supervised approach to enhance thermal-based depth estimation by leveraging pre-trained visual models initially designed for RGB data. In detail, we design a novel two-stage training strategy, incorporating Low-rank Adapters and Convolutional Adapters, which not only significantly improves accuracy and robustness but also enables impressive zero-shot generalization capabilities. Our method outperforms existing thermal-based depth estimation models, opening new possibilities for cross-modal applications in computer vision and robotics research. Ruoyu Fan, Wang Zhao 0001, Matthieu Lin, Qi Wang 0079, Yong-Jin Liu 0001, Wenping Wang 0001 |
ICRA | 4 |
| 2024 | Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 6 |
| 2024 | MISL: Multi-grained image-text semantic learning for text-guided image inpainting
Xingcai Wu, Kejun Zhao, Qianding Huang, Qi Wang 0079, Zhenguo Yang, Gefei Hao |
Pattern Recognit. | 4 |
| 2024 | SD-FSOD: Self-Distillation Paradigm via Distribution Calibration for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to detect novel targets with only a few instances of the associated samples. Although combinations of distillation techniques and meta-learning paradigms have been acknowledged as the primary strategies for FSOD tasks, the existing distillation methods exhibit inherent biases and sensitivity to novel class variability. A critical hurdle for FSOD distillation is the difficulty in ensuring appropriate knowledge learned from the teacher model during the fine-tuning stage. Furthermore, coarse distillation procedures risk misalignment between the learned and actual distributions. This misalignment could potentially negate the benefits of positive cases and impede the detector’s evolution. To address these deficiencies, we propose a novel self-distillation paradigm exclusively for the fine-tuning stage (SD-FSOD). Our methods integrate a Distribution Prototype Extractor (DPE) and Self-Distillation Memory (SDM), promoting feature distribution consistency during distillation. In detail, the DPE module reliably initializes the weights of the detector, ensuring a robust class distribution for the distillation process. Meanwhile, the SDM module utilizes decoupling techniques to divide the distillation tasks into two sub-task branches, allowing the student model to independently learn and share precise features through isolated distillation processes. The synergistic integration of feature calibration techniques and the continuous self-distillation paradigm distinctly enhances the fine-tuning process, which shows the superiority of the FSOD self-distillation methodologies. The extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our proposed approach produces significant improvements and achieves state-of-the-art (SOTA) performance. Qi Wang 0079, Kailin Xie, Liang Lei, Matthieu Lin, Tian Lv, Yong-Jin Liu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | FSNA: Few-Shot Object Detection via Neighborhood Information Adaption and All AttentionabstractFew-shot object detection (FSOD), a formidable task centered around developing inclusive models with annotated constrained samples, has attracted increasing interest in recent years. This discipline addresses unbalanced data distributions, which are particularly relevant to authentic scenarios. Although recent FSOD efforts have achieved considerable success in terms of localization, recognition remains a formidable obstacle. This stems from the fact that typical FSOD models evolve from general object detection frameworks predicated on extensive training data, and they underutilize and mine data information in scenarios with restricted samples, resulting in subpar performance. To address this deficiency, we introduce a groundbreaking methodology that is specifically tailored to overcome the inadequate sample challenge in FSOD tasks. Our approach incorporates a neighborhood information adaption (NIA) module that is designed to dynamically utilize information near the target, assisting in robustly performing object identification within the target domain. In addition, we propose an innovative attention mechanism called all attention, which not only encapsulates the dependencies of each position within a single feature map but also leverages correlations with other feature maps. This methodology culminates in more refined feature representations, which are particularly advantageous in situations with limited data. Comprehensive experiments conducted on the PASCAL VOC and COCO datasets illustrate that our technique achieves a substantial improvement with regard to addressing the FSOD task. Jinxiang Zhu, Qi Wang 0079, Weijian Ruan, Liang Lei, Gefei Hao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense CaptioningabstractThe dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods. Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | ALFPN: Adaptive Learning Feature Pyramid Network for Small Object DetectionabstractObject detection has become a crucial technology in intelligent vision systems, enabling automatic detection of target objects. While most detectors perform well on open datasets, they often struggle with small‐scale objects. This is due to the traditional top‐down feature fusion methods that weaken the semantic and location information of small objects, leading to poor classification performance. To address this issue, we propose a novel feature pyramid network, the adaptive learnable feature pyramid network (ALFPN). Our approach features an adaptive feature inspection that incorporates learnable fusion coefficients in the fusion of different levels of feature layers, aiding the network in learning features with less noise. In addition, we construct a context‐aligned supervisor that adjusts the feature maps fused at different levels to avoid scaling‐related offset effects. Our experiments demonstrate that our method achieves state‐of‐the‐art results and is highly robust for the small object detection on the TT‐100K, PASCAL VOC, and COCO datasets. These findings indicate that a model’s ability to extract discriminant features is positively correlated with its performance in detecting small objects. Qi Wang 0079, Weijian Ruan, Jingxiang Zhu, Liang Lei, Gefei Hao |
Int. J. Intell. Syst. | 2 |
| 2023 | GDENet: Graph Differential Equation Network for Traffic Flow PredictionabstractThe accurate prediction of traffic flow is paramount for the advancement of intelligent transportation systems. Despite this, current prediction models only account for either temporal or spatial features in isolation, without considering their interaction, impeding the model’s ability to express itself. In light of this, we propose the graph differential equations network (GDENet), an approach that can effectively mine spatiotemporal correlation. Specifically, we propose a spatiotemporal feature integrator (STFI), which alleviates the error caused by the deviation of the sampling distribution from the overall distribution. By incorporating temporal information into the model for training and combining it with spatial features, we thoroughly explore the spatiotemporal intrinsic association. When compared to state‐of‐the‐art methods, our proposed algorithm reduces memory consumption and elevates computational efficiency and the practical value. We conduct experiments with real‐world datasets, and our proposed model outperformed advanced prediction models. Yanming Miao, Xianghong Tang, Qi Wang 0079, Liya Yu |
Int. J. Intell. Syst. | 3 |
| 2023 | LCM-Captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text
Qi Wang 0079, Hongyu Deng, Zhenguo Yang, Yazhou Wang 0006, Gefei Hao |
Neural Networks | 1 |
| 2023 | Miper-MVS: Multi-scale iterative probability estimation with refinement for efficient multi-view stereo
Huizhou Zhou, Haoliang Zhao, Qi Wang 0079, Gefei Hao, Liang Lei |
Neural Networks | 3 |
| 2023 | AA-trans: Core attention aggregating transformer with information entropy selector for fine-grained visual classification
Qi Wang 0079, JianJun Wang, Hongyu Deng, Yazhou Wang 0006, Gefei Hao |
Pattern Recognit. | 1 |
| 2021 | Learning representation from multiple media domains for enhanced event discovery
Zhenguo Yang, Qing Li 0001, Haoran Xie 0001, Qi Wang 0079, Wenyin Liu |
Pattern Recognit. | 4 |
| 2020 | A novel feature representation: Aggregating convolution kernels for image retrieval
Qi Wang 0079, Jinxing Lai, Luc Claesen, Zhenguo Yang, Liang Lei, Wenyin Liu |
Neural Networks | 1 |
| 2020 | MetaSearch: Incremental Product Search via Deep Meta-LearningabstractWith the advancement of image processing and computer vision technology, content-based product search is applied in a wide variety of common tasks, such as online shopping, automatic checkout systems, and intelligent logistics. Given a product image as a query, existing product search systems mainly perform the retrieval process using predefined databases with fixed product categories. However, real-world applications often require inserting new categories or updating existing products in the product database. When using existing product search methods, the image feature extraction models must be retrained and database indexes must be rebuilt to accommodate the updated data, and these operations incur high costs for data annotation and training time. To this end, we propose a few-shot incremental product search framework with meta-learning, which requires very few annotated images and has a reasonable training time. In particular, our framework contains a multipooling-based product feature extractor that learns a discriminative representation for each product, and we also design a meta-learning-based feature adapter to guarantee the robustness of the few-shot features. Furthermore, when expanding new categories in batches during a product search, we reconstruct the few-shot features by using an incremental weight combiner to accommodate the incremental search task. Through extensive experiments, we demonstrate that the proposed framework achieves excellent performance for new products while still guaranteeing the high search accuracy of the base categories after gradually expanding new product categories without forgetting. Qi Wang 0079, Xinchen Liu, Wu Liu 0005, Anan Liu, Wenyin Liu, Tao Mei 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Improving cross-dimensional weighting pooling with multi-scale feature fusion for image retrieval
Qi Wang 0079, Jinxiang Lai, Zhenguo Yang, Kai Xu 0010, Peipei Kang, Wenyin Liu, Liang Lei |
Neurocomputing | 1 |
| 2018 | Beauty Product Image Retrieval Based on Multi-Feature Fusion and Feature AggregationabstractWe propose a beauty product image retrieval method based on multi-feature fusion and feature aggregation. The key idea is representing the image with the feature vector obtained by multi-feature fusion and feature aggregation. VGG16 and ResNet50 are chosen to extract image features, and Crow is adopted to perform deep feature aggregation. Benefited from the idea of transfer learning, we fine turn VGG16 on the Perfect-500K data set to improve the performance of image retrieval. The proposed method won the third price in Perfect Corp. Challenge 2018 with the best result 0.270676 mAP. We released our code on GitHub: https://github.com/wangqi12332155/ACMMM-beauty-AI-challenge. Qi Wang 0079, Jingxiang Lai, Kai Xu 0010, Wenyin Liu, Liang Lei |
ACM Multimedia | 1 |