Qi Wang 0079

dblp:19/1924-79 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-6269-0196ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WheatGOAT: Generalizable object-aware tracker via discriminative region semantic learning for wheat ear counting
Xingcai Wu, Yaoxi Li, Ziang Zou, Ya Yu, G. M. A. D. Sirishantha, A. S. A. Salgadoeb, Gefei Hao, Qi Wang 0079
Adv. Eng. Informatics9
2026 Quantifying expressive power in knowledge graph embeddings: An entropy-based metric framework
Panfeng Chen, Hui Li 0046, Qi Wang 0079, Xibin Wang
Expert Syst. Appl.3
2026 DHS-ViG: Dynamic hierarchical selective graph for comprehensive and robust feature perception
Chaojie Chen, Xingcai Wu, Yuanyuan Xiao, Peijia Yu, Qi Wang 0079
Neurocomputing5
2026 AdaptiveMamba: Comprehensive visual representation learning with adaptive semantic perception
Chaojie Chen, Xingcai Wu, Wu Liu 0005, Qi Wang 0079
Knowl. Based Syst.4
2025 A generalized CNN decision boundary theory based on non-Euclidean space
Yongjun Zhou, Qinglei Li, Yuanyuan Xiao, Qi Wang 0079
Appl. Intell.6
2025 Relation Semantic Guidance and Entity Position Location for Relation Extraction
abstract
Abstract Relation extraction is a research hot-spot in the field of natural language processing, and aims at structured knowledge acquirement. However, existing methods still grapple with the issue of entity overlapping, where they treat relation types as inconsequential labels, overlooking the fact that relation type has a great influence on entity type hindering the performance of these models from further improving. Furthermore, current models are inadequate in handling the fine-grained aspect of entity positioning, which leads to ambiguity in entity boundary localization and uncertainty in relation inference, directly. In response to this challenge, a relation extraction model is proposed, which is guided by relational semantic cues and focused on entity boundary localization. The model uses an attention mechanism to align relation semantics with sentence information, so as to obtain the most relevant semantic expression to the target relation instance. It then incorporates an entity locator to harness additional positional features, thereby, enhancing the capability of the model to pinpoint entity start and end tags. Consequently, this approach effectively alleviates the problem of entity overlapping. Extensive experiments are conducted on the widely used datasets NYT and WebNLG. The experimental results show that the proposed model outperforms the baseline ones in F1 scores of the two datasets, and the improvement margin is up to 5.50% and 2.80%, respectively.
Panfeng Chen, Hui Li 0046, Xibin Wang, Aihua Yu, Xingzhi Deng, Qi Wang 0079
Data Sci. Eng.8
2025 EMGE: Entities and Mentions Gradual Enhancement with semantics and connection modelling for document-level relation extraction
Panfeng Chen, Qi Wang 0079, Hui Li 0046, Xibin Wang, Aihua Yu, Xingzhi Deng
Knowl. Based Syst.3
2025 DRC: Discrete Representation Classifier With Salient Features via Fixed-Prototype
abstract
Image classification models including convolutional neural networks (CNN) and vision transformers (ViT) commonly employ a fully connected (FC) layer as the classifier. However, the fully connected nature of FC brings large amounts of weight parameters, limits the efficiency of inference, tends to over-fit the training data, and struggles to learn distinct class weights. To solve these problems, we propose a discrete representation classifier (DRC), a generic parameter-free classifier that offers efficiency, robustness, and more discriminative categorization. Specifically, the DRC discards numerous unimportant features and focuses solely on the salient features which are reinforced during training and presented in short discrete form during inference. Unlike the way of learning pseudo-prototypes (weights) from data laden with complex patterns and noises in FC, the DRC introducing discriminative fixed-prototypes which are almost uniformly distributed across the high-dimensional feature space, thus helps the model to learn more distinct boundaries between categories. Further leveraging the advantage of DRC’s focus on salient features, we propose Salient-CAM, which is able to locate the most important region in image without the need for weighting feature maps. The experiments demonstrate that simply replacing the model’s classifier from FC to DRC can lead to a significant acceleration in the whole model’s inference and a more robust classification. Additionally, the proposed Salient-CAM exhibits excellent object localization ability in complex natural scenes.
Qinglei Li, Qi Wang 0079, Yongbin Qin, Xingcai Wu, Shiming Chen 0002, Wu Liu 0005, Yong-Jin Liu 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation Model
abstract
Depth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains challenging in robust generalization due to limited data resources and spectral differences between thermal and RGB images. In this paper, we present a self-supervised approach to enhance thermal-based depth estimation by leveraging pre-trained visual models initially designed for RGB data. In detail, we design a novel two-stage training strategy, incorporating Low-rank Adapters and Convolutional Adapters, which not only significantly improves accuracy and robustness but also enables impressive zero-shot generalization capabilities. Our method outperforms existing thermal-based depth estimation models, opening new possibilities for cross-modal applications in computer vision and robotics research.
Ruoyu Fan, Wang Zhao 0001, Matthieu Lin, Qi Wang 0079, Yong-Jin Liu 0001, Wenping Wang 0001
ICRA4
2024 Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001
Comput. Aided Geom. Des.6
2024 MISL: Multi-grained image-text semantic learning for text-guided image inpainting
Xingcai Wu, Kejun Zhao, Qianding Huang, Qi Wang 0079, Zhenguo Yang, Gefei Hao
Pattern Recognit.4
2024 SD-FSOD: Self-Distillation Paradigm via Distribution Calibration for Few-Shot Object Detection
abstract
Few-shot object detection (FSOD) aims to detect novel targets with only a few instances of the associated samples. Although combinations of distillation techniques and meta-learning paradigms have been acknowledged as the primary strategies for FSOD tasks, the existing distillation methods exhibit inherent biases and sensitivity to novel class variability. A critical hurdle for FSOD distillation is the difficulty in ensuring appropriate knowledge learned from the teacher model during the fine-tuning stage. Furthermore, coarse distillation procedures risk misalignment between the learned and actual distributions. This misalignment could potentially negate the benefits of positive cases and impede the detector’s evolution. To address these deficiencies, we propose a novel self-distillation paradigm exclusively for the fine-tuning stage (SD-FSOD). Our methods integrate a Distribution Prototype Extractor (DPE) and Self-Distillation Memory (SDM), promoting feature distribution consistency during distillation. In detail, the DPE module reliably initializes the weights of the detector, ensuring a robust class distribution for the distillation process. Meanwhile, the SDM module utilizes decoupling techniques to divide the distillation tasks into two sub-task branches, allowing the student model to independently learn and share precise features through isolated distillation processes. The synergistic integration of feature calibration techniques and the continuous self-distillation paradigm distinctly enhances the fine-tuning process, which shows the superiority of the FSOD self-distillation methodologies. The extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our proposed approach produces significant improvements and achieves state-of-the-art (SOTA) performance.
Qi Wang 0079, Kailin Xie, Liang Lei, Matthieu Lin, Tian Lv, Yong-Jin Liu 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 FSNA: Few-Shot Object Detection via Neighborhood Information Adaption and All Attention
abstract
Few-shot object detection (FSOD), a formidable task centered around developing inclusive models with annotated constrained samples, has attracted increasing interest in recent years. This discipline addresses unbalanced data distributions, which are particularly relevant to authentic scenarios. Although recent FSOD efforts have achieved considerable success in terms of localization, recognition remains a formidable obstacle. This stems from the fact that typical FSOD models evolve from general object detection frameworks predicated on extensive training data, and they underutilize and mine data information in scenarios with restricted samples, resulting in subpar performance. To address this deficiency, we introduce a groundbreaking methodology that is specifically tailored to overcome the inadequate sample challenge in FSOD tasks. Our approach incorporates a neighborhood information adaption (NIA) module that is designed to dynamically utilize information near the target, assisting in robustly performing object identification within the target domain. In addition, we propose an innovative attention mechanism called all attention, which not only encapsulates the dependencies of each position within a single feature map but also leverages correlations with other feature maps. This methodology culminates in more refined feature representations, which are particularly advantageous in situations with limited data. Comprehensive experiments conducted on the PASCAL VOC and COCO datasets illustrate that our technique achieves a substantial improvement with regard to addressing the FSOD task.
Jinxiang Zhu, Qi Wang 0079, Weijian Ruan, Liang Lei, Gefei Hao
IEEE Trans. Circuits Syst. Video Technol.2
2024 CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense Captioning
abstract
The dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods.
Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001
IEEE Trans. Multim.3
2023 ALFPN: Adaptive Learning Feature Pyramid Network for Small Object Detection
abstract
Object detection has become a crucial technology in intelligent vision systems, enabling automatic detection of target objects. While most detectors perform well on open datasets, they often struggle with small‐scale objects. This is due to the traditional top‐down feature fusion methods that weaken the semantic and location information of small objects, leading to poor classification performance. To address this issue, we propose a novel feature pyramid network, the adaptive learnable feature pyramid network (ALFPN). Our approach features an adaptive feature inspection that incorporates learnable fusion coefficients in the fusion of different levels of feature layers, aiding the network in learning features with less noise. In addition, we construct a context‐aligned supervisor that adjusts the feature maps fused at different levels to avoid scaling‐related offset effects. Our experiments demonstrate that our method achieves state‐of‐the‐art results and is highly robust for the small object detection on the TT‐100K, PASCAL VOC, and COCO datasets. These findings indicate that a model’s ability to extract discriminant features is positively correlated with its performance in detecting small objects.
Qi Wang 0079, Weijian Ruan, Jingxiang Zhu, Liang Lei, Gefei Hao
Int. J. Intell. Syst.2
2023 GDENet: Graph Differential Equation Network for Traffic Flow Prediction
abstract
The accurate prediction of traffic flow is paramount for the advancement of intelligent transportation systems. Despite this, current prediction models only account for either temporal or spatial features in isolation, without considering their interaction, impeding the model’s ability to express itself. In light of this, we propose the graph differential equations network (GDENet), an approach that can effectively mine spatiotemporal correlation. Specifically, we propose a spatiotemporal feature integrator (STFI), which alleviates the error caused by the deviation of the sampling distribution from the overall distribution. By incorporating temporal information into the model for training and combining it with spatial features, we thoroughly explore the spatiotemporal intrinsic association. When compared to state‐of‐the‐art methods, our proposed algorithm reduces memory consumption and elevates computational efficiency and the practical value. We conduct experiments with real‐world datasets, and our proposed model outperformed advanced prediction models.
Yanming Miao, Xianghong Tang, Qi Wang 0079, Liya Yu
Int. J. Intell. Syst.3
2023 LCM-Captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text
Qi Wang 0079, Hongyu Deng, Zhenguo Yang, Yazhou Wang 0006, Gefei Hao
Neural Networks1
2023 Miper-MVS: Multi-scale iterative probability estimation with refinement for efficient multi-view stereo
Huizhou Zhou, Haoliang Zhao, Qi Wang 0079, Gefei Hao, Liang Lei
Neural Networks3
2023 AA-trans: Core attention aggregating transformer with information entropy selector for fine-grained visual classification
Qi Wang 0079, JianJun Wang, Hongyu Deng, Yazhou Wang 0006, Gefei Hao
Pattern Recognit.1
2021 Learning representation from multiple media domains for enhanced event discovery
Zhenguo Yang, Qing Li 0001, Haoran Xie 0001, Qi Wang 0079, Wenyin Liu
Pattern Recognit.4
2020 A novel feature representation: Aggregating convolution kernels for image retrieval
Qi Wang 0079, Jinxing Lai, Luc Claesen, Zhenguo Yang, Liang Lei, Wenyin Liu
Neural Networks1
2020 MetaSearch: Incremental Product Search via Deep Meta-Learning
abstract
With the advancement of image processing and computer vision technology, content-based product search is applied in a wide variety of common tasks, such as online shopping, automatic checkout systems, and intelligent logistics. Given a product image as a query, existing product search systems mainly perform the retrieval process using predefined databases with fixed product categories. However, real-world applications often require inserting new categories or updating existing products in the product database. When using existing product search methods, the image feature extraction models must be retrained and database indexes must be rebuilt to accommodate the updated data, and these operations incur high costs for data annotation and training time. To this end, we propose a few-shot incremental product search framework with meta-learning, which requires very few annotated images and has a reasonable training time. In particular, our framework contains a multipooling-based product feature extractor that learns a discriminative representation for each product, and we also design a meta-learning-based feature adapter to guarantee the robustness of the few-shot features. Furthermore, when expanding new categories in batches during a product search, we reconstruct the few-shot features by using an incremental weight combiner to accommodate the incremental search task. Through extensive experiments, we demonstrate that the proposed framework achieves excellent performance for new products while still guaranteeing the high search accuracy of the base categories after gradually expanding new product categories without forgetting.
Qi Wang 0079, Xinchen Liu, Wu Liu 0005, Anan Liu, Wenyin Liu, Tao Mei 0001
IEEE Trans. Image Process.1
2019 Improving cross-dimensional weighting pooling with multi-scale feature fusion for image retrieval
Qi Wang 0079, Jinxiang Lai, Zhenguo Yang, Kai Xu 0010, Peipei Kang, Wenyin Liu, Liang Lei
Neurocomputing1
2018 Beauty Product Image Retrieval Based on Multi-Feature Fusion and Feature Aggregation
abstract
We propose a beauty product image retrieval method based on multi-feature fusion and feature aggregation. The key idea is representing the image with the feature vector obtained by multi-feature fusion and feature aggregation. VGG16 and ResNet50 are chosen to extract image features, and Crow is adopted to perform deep feature aggregation. Benefited from the idea of transfer learning, we fine turn VGG16 on the Perfect-500K data set to improve the performance of image retrieval. The proposed method won the third price in Perfect Corp. Challenge 2018 with the best result 0.270676 mAP. We released our code on GitHub: https://github.com/wangqi12332155/ACMMM-beauty-AI-challenge.
Qi Wang 0079, Jingxiang Lai, Kai Xu 0010, Wenyin Liu, Liang Lei
ACM Multimedia1