VLDB 2026 Research / reviewers in the wild / expert
Bingbing Zhang 0001
dblp:131/4775-1
· DBLP profile ↗
18ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-4734-4164ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M 2 -CLIP++: Dual-branch high-order vision-language adaptation with dynamic semantic prompting for efficient video recognition
Bingbing Zhang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2026 | QHSP-Net: query-aware higher-order statistical pooling network for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li |
Multim. Syst. | 3 |
| 2026 | An Asynchronous Intermittent Control Methodology for Cyber-Physical Systems Under Dynamic Actuator Faults
Ruoqi Li, Bingbing Zhang 0001, Yang Yang 0052, Qi-He Shan, Lei Liu 0006 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Dual-Prompt Learning with Cross-Modal Decoders for Few-Shot Whole Slide Image ClassificationabstractFew-shot learning offers a promising solution for computational pathology by alleviating the reliance on large an-notated datasets, but faces challenges from the high redundancy in whole slide images and underutilized cross-modal knowledge. Existing methods typically use foundation models only for pre-liminary feature extraction while employing fixed or single-level prompts that lack multi-scale pathological representation. To address these limitations, we propose a Hierarchical Vision-Text Prompt (H- VTP) framework that enables multi-level cross-modal interaction through GPT-4 generated Local Instance Prompts for patch-level morphological details and Global Semantic Prompts for slide-level diagnostic context. A dual-branch decoding mech-anism with Text-Guided-Patch Decoder and Patch-Augmented-Text Decoder facilitates closed-loop vision-text fusion, while a parameter-efficient adaptation strategy trains only lightweight prompts and adapters. Extensive experiments on three cancer subtype datasets demonstrate the superiority of H-VTP in few-shot WSI classification, confirming its effectiveness for clinical applications. Bingbing Zhang 0001, Wen Zhu, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 2 |
| 2025 | Cross Attention Guided Multimodal Network for Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 1 |
| 2025 | High-Order Multimodal Multi-task Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 1 |
| 2025 | Few-Shot Action Recognition Based on Visual-Language Prototype Hierarchical Temporal Enhancement
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 1 |
| 2025 | Instance-aware context with mutually guided vision-language attention for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li |
Appl. Intell. | 3 |
| 2025 | A2 M2-Net: Adaptively Aligned Multi-scale Moment for Few-Shot Action Recognition
Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li |
Int. J. Comput. Vis. | 3 |
| 2024 | PM2: A New Prompting Multi-modal Model Paradigm for Few-shot Medical Image ClassificationabstractFew-shot learning has become a key technical solution for addressing the challenges of limited data and difficult annotation acquisition in medical image classification. However, relying solely on a single image modality proves inadequate for capture conceptual categories. This paper proposes a novel medical image classification paradigm based on a multi-modal foundation model, called PM2. In addition to the image modality, PM2introduces supplementary text input (prompt) to further describe images or conceptual categories and facilitate cross-modal few-shot learning. We empirically studied five different prompting schemes under this new paradigm. Furthermore, linear probing in multi-modal models only takes class token as input, ignoring the rich statistical data contained in high-level visual tokens. Therefore, we alternately perform linear classification on the feature distributions of visual tokens and class token. To effectively extract statistical information, we use global covariance pool with efficient matrix power normalization to aggregate the visual tokens. We then combine two classification heads: one for handling image class token and prompt representations encoded by the text encoder, and the other for classifying the feature distributions of visual tokens. Experiments on two medical datasets demonstrate that regardless of the prompting scheme, our method PM2outperforms its counterparts, achieving state-of-the-art performance. Zhenwei Wang 0005, Qiule Sun, Bingbing Zhang 0001, Weijian Su, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 3 |
| 2024 | Dynamic Temporal Shift Feature Enhancement for Few-Shot Action Recognition
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (10) | 2 |
| 2024 | Visual-guided hierarchical iterative fusion for multi-modal video action recognition
Bingbing Zhang 0001, Jianxin Zhang 0001, Qiule Sun, Qiang Zhang 0008 |
Pattern Recognit. Lett. | 1 |
| 2023 | GSoANet: Group Second-Order Aggregation Network for Video Action Recognition
Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Bin Liu 0040, Qiang Zhang 0008 |
Neural Process. Lett. | 3 |
| 2022 | High-order Correlation Network for Video RecognitionabstractHow to model global video representation is an important research content of video recognition. Among current convolutional neural network(CNN) based methods, only using first-order representations (i.e. global average pooling) has limitations in capturing spatiotemporal features of videos. Recent studies have shown that high-order statistics are more suitable to model complex feature distributions. To better characterize the spatiotemporal structure for video recognition, we propose a novel High-order Correlation Network (HoCNet) in this work, the core of which is to explore high-order video representations through correlation computation and covariance pooling. HoC-Net leverages the correlation module to obtain complex temporal dynamic information of frames via computing dot product of features in the fixed sliding window of two adjacent frames. As an approximate high-order calculation, the correlation module can be inserted into any stage of the deep network to model high-order representations in various spatial resolutions. Additionally, a robust high-order pooling module, i.e., iterative matrix square root normalization of covariance pooling (iSQRT-COV), is also introduced at the end of the network, and this further boosts modeling complex spatiotemporal distributions of video features. Experiments conducted on four widely used video benchmarks demonstrate the effectiveness of HoCNet, which achieves the comparable performance with the state-of-the-art models. Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Qiang Zhang 0008 |
IJCNN | 3 |
| 2022 | Temporal grafter network: Rethinking LSTM for effective video recognition
Bingbing Zhang 0001, Qilong Wang 0001, Zilin Gao, Ruiren Zeng, Peihua Li |
Neurocomputing | 1 |
| 2021 | Temporal-attentive Covariance Pooling Networks for Video RecognitionabstractFor video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to capture complex dynamics of videos. For image recognition task, there exist evidences showing that covariance pooling has stronger representation ability than GAP. Unfortunately, such plain covariance pooling used in image recognition is an orderless representative, which cannot model spatio-temporal structure inherent in videos. Therefore, this paper proposes a Temporal-attentive Covariance Pooling (TCP), inserted at the end of deep architectures, to produce powerful video representations. Specifically, our TCP first develops a temporal attention module to adaptively calibrate spatio-temporal features for the succeeding covariance pooling, approximatively producing attentive covariance representations. Then, a temporal covariance pooling performs temporal pooling of the attentive covariance representations to characterize both intra-frame correlations and inter-frame cross-correlations of the calibrated features. As such, the proposed TCP can capture complex temporal dynamics. Finally, a fast matrix power normalization is introduced to exploit geometry of covariance representations. Note that our TCP is model-agnostic and can be flexibly integrated into any video architectures, resulting in TCPNet for effective video recognition. The extensive experiments on six benchmarks (e.g., Kinetics, Something-Something V1 and Charades) using various video architectures show our TCPNet is clearly superior to its counterparts, while having strong generalization ability. The source code is publicly available. Zilin Gao, Qilong Wang 0001, Bingbing Zhang 0001, Qinghua Hu, Peihua Li |
NeurIPS | 3 |
| 2020 | Locality-constrained affine subspace coding for image classification and retrieval
Bingbing Zhang 0001, Qilong Wang 0001, Xiaoxiao Lu, Fasheng Wang, Peihua Li |
Pattern Recognit. | 1 |
| 2014 | A personalized ellipsoid modeling method and matching error analysis for the artificial femoral head design
Bin Liu 0040, Shungang Hua, Zhaoliang Liu, Bingbing Zhang 0001, Zongge Yue |
Comput. Aided Des. | 6 |