VLDB 2026 Research / reviewers in the wild / expert
Wei Zhou 0042
dblp:69/5011-42
· DBLP profile ↗
18ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-9237-7205ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omniscient bottom-up double-stream symmetric network for image captioning
Jianchao Li, Wei Zhou 0042, Kai Wang 0033, Haifeng Hu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | UniPT: A Unified Representation Pre-training for Multi-dataset 3D Object Detection
Zhijie Zheng 0003, Kang Lin, Wei Zhou 0042, Zhongyuan Qiu, Yali Zhao, Huan Qin, Dihu Chen |
PRICAI (5) | 5 |
| 2025 | From grids to pseudo-regions: Dynamic memory augmented image captioning with dual relation transformer
Wei Zhou 0042, Weitao Jiang, Zhijie Zheng 0003, Jianchao Li, Haifeng Hu 0001 |
Expert Syst. Appl. | 1 |
| 2025 | From multi-scale grids to dynamic regions: Dual-relation enhanced transformer for image captioning
Wei Zhou 0042, Chuanle Song, Dihu Chen, Haifeng Hu 0001, Chun Shan |
Knowl. Based Syst. | 1 |
| 2025 | DRTN: Dual Relation Transformer Network with feature erasure and contrastive learning for multi-label image classification
Wei Zhou 0042, Kang Lin, Zhijie Zheng 0003, Dihu Chen, Haifeng Hu 0001 |
Neural Networks | 1 |
| 2025 | Residual Quotient Learning for Zero-Reference Low-Light Image EnhancementabstractRecently, neural networks have become the dominant approach to low-light image enhancement (LLIE), with at least one-third of them adopting a Retinex-related architecture. However, through in-depth analysis, we contend that this most widely accepted LLIE structure is suboptimal, particularly when addressing the non-uniform illumination commonly observed in natural images. In this paper, we present a novel variant learning framework, termed residual quotient learning, to substantially alleviate this issue. Instead of following the existing Retinex-related decomposition-enhancement-reconstruction process, our basic idea is to explicitly reformulate the light enhancement task as adaptively predicting the latent quotient with reference to the original low-light input using a residual learning fashion. By leveraging the proposed residual quotient learning, we develop a lightweight yet effective network called ResQ-Net. This network features enhanced non-uniform illumination modeling capabilities, making it more suitable for real-world LLIE tasks. Moreover, due to its well-designed structure and reference-free loss function, ResQ-Net is flexible in training as it allows for zero-reference optimization, which further enhances the generalization and adaptability of our entire framework. Extensive experiments on various benchmark datasets demonstrate the merits and effectiveness of the proposed residual quotient learning, and our trained ResQ-Net outperforms state-of-the-art methods both qualitatively and quantitatively. Furthermore, a practical application in dark face detection is explored, and the preliminary results confirm the potential and feasibility of our method in real-world scenarios. Linfeng Fei, Huanjie Tao, Yaocong Hu, Wei Zhou 0042, Jiun Tian Hoe, Weipeng Hu, Yap-Peng Tan |
IEEE Trans. Image Process. | 5 |
| 2025 | Snippet-Inter Difference Attention Network for Weakly-Supervised Temporal Action LocalizationabstractThe purpose of weakly-supervised temporal action localization (WTAL) task is to simultaneously classify and localize action instances in untrimmed videos with only video-level labels. Previous works fail to extract multi-scale temporal features to identify action instances with different durations, and they do not fully use the temporal cues of action video to learn discriminative features. In addition, the classifiers trained by current methods usually focus on easy-to-distinguish snippets while ignoring other semantically ambiguous features, which leads to incomplete and over-complete localization. To address these issues, we introduce a new Snippet-inter Difference Attention Network (SDANet) for WTAL, which can be trained end-to-end. Specifically, our model presents three modules, with primary contributions lying in the snippet-inter difference attention (SDA) module and potential feature mining (PFM) module. Firstly, we construct a simple multi-scale temporal feature fusion (MTFF) module to generate multi-scale temporal feature representation, so as to help the model better detect short action instances. Secondly, we consider the temporal cues of video features and design SDA module based on the Transformer to capture global discriminative features for each modality based on multi-scale features. It calculates the differences between temporal neighbor snippets in each modality to explore salient-difference features, and then utilizes them to guide correlation modeling. Thirdly, after learning discriminative features, we devise PFM module to excavate potential action and background snippets from ambiguous features. By contrastive learning, potential actions are forced closer to discriminative actions and away from the background, thereby learning more accurate action boundaries. Finally, two losses (i.e., similarity loss and reconstruction loss) are further developed to constrain the consistency between two modalities and help the model retain original feature information for better localization results. Extensive experiments show that our model achieves better performance against current WTAL methods on three datasets, i.e., THUMOS14, ActivityNet1.2 and ActivityNet1.3. Wei Zhou 0042, Kang Lin, Weipeng Hu, Haifeng Hu 0001, Yap-Peng Tan |
IEEE Trans. Multim. | 1 |
| 2025 | Temporal and Semantic Correlation Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-Supervised Temporal Action Localization (WTAL) aims to identify the temporal boundaries and classify actions in untrimmed videos using only video-level labels during training. Despite recent progress, many existing approaches primarily follow a localization-by-classification pipeline, treating snippets as independent instances and thus exploiting only limited contextual information. Besides, these methods struggle to capture multi-scale temporal information and neglect both the internal temporal structures within videos and the semantic consistency between videos, resulting in misclassification and inaccurate localization. To address these limitations, we introduce a novel Temporal and Semantic Correlation Network (TSC-Net) for WTAL task, which can be trained end-to-end. First, we propose a Multi-Scale Features Integration Pyramid (MFIP) module to integrate multi-scale temporal features, effectively addressing the challenge of missed detections caused by short action durations. Furthermore, we design a Temporal Correlation Enhancement (TCE) branch to enhance segment correlations by video-level temporal structures to improve the completeness of action localization. Finally, a Dataset-Wide Semantic Awareness (DSA) branch is designed to construct and propagate a dataset-level action semantics bank, enhancing the model’s awareness of semantic consistency in actions. Extensive experiments show that TSC-Net outperforms most existing WTAL methods, achieving an average mAP of 46.3% on the THUMOS-14 dataset and 26.5% on the ActivityNet1.2 dataset. Detailed ablation studies further confirm the effectiveness of each component in our model. The code and models are publicly available at https://github.com/linkang-els/TSC-Net-main . Kang Lin, Wei Zhou 0042, Zhijie Zheng 0003, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | HIN: Hierarchical Interaction Network for Image CaptioningabstractThe purpose of the image captioning task is to understand the content of an image and generate corresponding descriptive text. Traditional approaches to image captioning typically generate descriptive text by extracting different types of visual features from an image and performing feature interactions. However, these methods often fail to fully exploit the interactions between different types of visual features, leading to suboptimal feature integration. To address this limitation, we propose a novel Hierarchical Interaction Network (HIN) , designed to continuously extract and interact with different types of visual features to perform more effective multilevel feature interactions. Our HIN consists of three key modules: firstly, we design the Cross-Type Feature Alignment (CTFA) encoder, which aligns different types of visual features by three global features, so that the subsequent modules can effectively carry out the Hierarchical Interaction (HI) ; secondly, the HI module, which utilizes different types of multilevel features output from the encoder to carry out feature interactions and information mining, so as to generate fully mined multilevel features. The Bottom-up Gated Attention Fusion (BGAF) decoder is finally designed to perform the multilevel decoding of the features mined by our HI module, further enhancing the feature interaction capabilities of our HIN. Moreover, additional experiments on the MS-COCO dataset show that our model achieves new state-of-the-art performance. All codes are available at https://github.com/songchuanle-1/HIN . Chuanle Song, Wei Zhou 0042, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Triple-Stream Commonsense Circulation Transformer Network for Image Captioning
Jianchao Li, Wei Zhou 0042, Kai Wang 0033, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2024 | DATran: Dual Attention Transformer for Multi-Label Image ClassificationabstractMulti-label image classification is a fundamental yet challenging task, which aims to predict the labels associated with a given image. Most of previous methods directly exploit the high-level features from the last layer of convolutional neural network for classification. However, these methods cannot obtain global features due to the limited size of convolutional kernels, and they fail to extract multi-scale features to effectively recognize small-scale objects in the images. Recent studies exploit the graph convolution network to model the label correlations for boosting the classification performance. Despite substantial progress, these methods rely on manually pre-defined graph structures. Besides, they ignore the associations between semantic labels and image regions, and do not fully explore the spatial context of images. To address above issues, we propose a novel Dual Attention Transformer (DATran) model, which adopts a dual-stream architecture that simultaneously learns spatial and channel correlations from multi-label images. Firstly, in order to solve the problem that current methods are difficult to recognize small-size objects, we develop a new multi-scale feature fusion (MSFF) module to generate multi-scale feature representation by jointly integrating both high-level semantics and low-level details. Secondly, we design a prior-enhanced spatial attention (PSA) module to learn the long-range correlation between objects from different spatial positions in images to enhance the model performance. Thirdly, we devise a prior-enhanced channel attention (PCA) module to capture the inter-dependencies between different channel maps, thus effectively improving the correlation between semantic categories. It is worth noting that PSA module and PCA module complement and promote each other to further augment the feature representations. Finally, the outputs of these two attention modules are fused to obtain the final features for classification. Performance evaluation experiments are conducted on MS-COCO 2014, PASCAL VOC 2007 and VG-500 datasets, demonstrating that DATran model achieves better performance than current state-of-the-art models. Wei Zhou 0042, Zhijie Zheng 0003, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Mining Semantic Information With Dual Relation Graph Network for Multi-Label Image ClassificationabstractThe purpose of multi-label image classification is to assign multiple labels for multiple objects presented in one image. Recent research efforts exploit graph convolution network (GCN) to learn the label co-occurrence dependencies for enhancing the semantic representation. Although these methods have achieved promising results, they can not capture the intrinsic correlation between objects in images and do not consider the inter-channel relationship. In addition, the previous methods treat each single image independently and fail to explore the relationship between different images. To address the above challenges, we propose a novelDualRelationGraphNetwork (DRGN) model, which adopts a double branch structure to excavate rich semantic information from intra-image and cross-image simultaneously. Specifically, we first develop an intra-image channel-relation mining (ICM) module to mine the inter-channel relationship in features while learning the importance of different channels. Secondly, we design a new GCN-based intra-image spatial-relation exploring (ISE) module to capture the correlation between objects in individual image. Notably, ISE module and ICM module can complement and promote each other from the spatial and channel dimensions of images to improve the correlation between objects in individual image. Thirdly, we propose a novel GCN-based cross-image semantic learning (CSL) module to learn the semantic relationship between different images in the mini-batch. Through graph reasoning, our CSL module can iteratively refine input image features by acquiring common semantic information from other images in the mini-batch. Extensive experiments on the MS-COCO 2014, PASCAL VOC 2007, and VG-500 datasets demonstrate that the proposed DRGN model outperforms current state-of-the-art methods. Wei Zhou 0042, Weitao Jiang, Dihu Chen, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Feature learning network with transformer for multi-label image classification
Wei Zhou 0042, Peng Dou, Haifeng Hu 0001, Zhijie Zheng 0003 |
Pattern Recognit. | 1 |
| 2023 | Attention-Augmented Memory Network for Image Multi-Label ClassificationabstractThe purpose of image multi-label classification is to predict all the object categories presented in an image. Some recent works exploit graph convolution network to capture the correlation between labels. Although promising results have been reported, these methods cannot learn salient object features in the images and ignore the correlation between channel feature maps. In addition, the current researches only learn the feature information within individual input image, but fail to mine the contextual information of various categories from the dataset to enhance the input feature representation. To address these issues, we propose an A ttention- A ugmented M emory N etwork ( AAMN ) model for the image multi-label classification task. Specifically, we first propose a novel categorical memory module to excavate the contextual information of various categories from the dataset to augment the current input feature. Secondly, we design a new channel-relation exploration module to capture the inter-channel relationship of features, so as to enhance the correlation between objects in the images. Thirdly, we develop a spatial-relation enhancement module to model second-order statistics of features and capture long-range dependencies between pixels in feature maps, so as to learn salient object features. Experimental results on standard benchmarks, including MS-COCO 2014, PASCAL VOC 2007, and VG-500, demonstrate the effectiveness and superiority of AAMN model, which outperforms current state-of-the-art methods. Wei Zhou 0042, Yanke Hou, Dihu Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Aligning Image Semantics and Label Concepts for Image Multi-Label ClassificationabstractImage multi-label classification task is mainly to correctly predict multiple object categories in the images. To capture the correlation between labels, graph convolution network based methods have to manually count the label co-occurrence probability from training data to construct a pre-defined graph as the input of graph network, which is inflexible and may degrade model generalizability. Moreover, most of the current methods cannot effectively align the learned salient object features with the label concepts, so that the predicted results of model may not be consistent with the image content. Therefore, how to learn the salient semantic features of images and capture the correlation between labels, and then effectively align them is one of the key to improve the performance of image multi-label classification task. To this end, we propose a novel image multi-label classification framework which aims to align I mage S emantics with L abel C oncepts ( ISLC ). Specifically, we propose a residual encoder to learn salient object features in the images, and exploit the self-attention layer in aligned decoder to automatically capture the correlation between labels. Then, we leverage the cross-attention layers in aligned decoder to align image semantic features with label concepts, so as to make the labels predicted by model more consistent with image content. Finally, the output features of the last layer of residual encoder and aligned decoder are fused to obtain the final output feature for classification. The proposed ISLC model achieves good performance on various prevalent multi-label image datasets such as MS-COCO 2014, PASCAL VOC 2007, VG-500, and NUS-WIDE with 87.2%, 96.9%, 39.4%, and 64.2%, respectively. Wei Zhou 0042, Zhiwu Xia, Peng Dou, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Double Attention Based on Graph Attention Network for Image Multi-Label ClassificationabstractThe task of image multi-label classification is to accurately recognize multiple objects in an input image. Most of the recent works need to leverage the label co-occurrence matrix counted from training data to construct the graph structure, which are inflexible and may degrade model generalizability. In addition, these methods fail to capture the semantic correlation between the channel feature maps to further improve model performance. To address these issues, we propose DA-GAT (a D ouble A ttention framework based on the G raph A ttention ne T work) to effectively learn the correlation between labels from training data. First, we devise a new channel attention mechanism to enhance the semantic correlation between channel feature maps, so as to implicitly capture the correlation between labels. Second, we propose a new label attention mechanism to avoid the adverse impact of a manually constructed label co-occurrence matrix. It only needs to leverage the label embedding as the input of network, then automatically constructs the label relation matrix to explicitly establish the correlation between labels. Finally, we effectively fuse the output of these two attention mechanisms to further improve model performance. Extensive experiments are conducted on three public multi-label classification benchmarks. Our DA-GAT model achieves mean average precision of 87.1%, 96.6%, and 64.3% on MS-COCO 2014, PASCAL VOC 2007, and NUS-WIDE, respectively, and obviously outperforms other existing state-of-the-art methods. In addition, visual analysis experiments demonstrate that each attention mechanism can capture the correlation between labels well and significantly promote the model performance. Wei Zhou 0042, Zhiwu Xia, Peng Dou, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Double-Stream Position Learning Transformer Network for Image CaptioningabstractImage captioning has made significant achievement through developing feature extractor and model architecture. Recently, the image region features extracted by object detector prevail in most existing models. However, region features are criticized for the lacking of background and full contextual information. This problem can be remedied by providing some complementary visual information from patch features. In this paper, we propose a Double-Stream Position Learning Transformer Network (DSPLTN) which exploits the advantages of region features and patch features. Specifically, the region-stream encoder utilizes a Transformer encoder with Relative Position Learning (RPL) module to enhance the representations of region features through modeling the relationships between regions and positions respectively. As for the patch-stream encoder, we introduce convolutional neural network into the vanilla Transformer encoder and propose a novel Convolutional Position Learning (CPL) module to encode the position relationships between patches. CPL improves the ability of relationship modeling by combining the position and visual content of patches. Incorporating CPL into the Transformer encoder can synthesize the benefits of convolution in local relation modeling and self-attention in global feature fusion, thereby compensating for the information loss caused by the flattening operation of 2D feature maps to 1D patches. Furthermore, an Adaptive Fusion Attention (AFA) mechanism is proposed to balance the contribution of enhanced region and patch features. Extensive experiments on MSCOCO demonstrate the effectiveness of the double-stream encoder and CPL, and show the superior performance of DSPLTN. Weitao Jiang, Wei Zhou 0042, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Feature Matching Network for Weakly-Supervised Temporal Action Localization
Peng Dou, Wei Zhou 0042, Zhongke Liao, Haifeng Hu 0001 |
PRCV (4) | 2 |