VLDB 2026 Research / reviewers in the wild / expert
Hongshen Zhao
dblp:373/4477
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0005-6314-0152ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Deep learning architectures and training · 23% Efficient and distributed learning · 23% Vision and language · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.9 | 1 | 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Mamba-Adaptor: State Space Model Adaptor for Visual Recognition · CVPR 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.9 | 1 | 2025 | Mamba-Adaptor: State Space Model Adaptor for Visual Recognition · CVPR 2025 |
Computer vision › Vision and language
visual grounding |
0.9 | 1 | 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.3 | 1 | 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
proposal-based framework · 0.9multi-granularity discrimination · 0.9memory augmentation · 0.9dilated convolution · 0.9contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mamba-Adaptor: State Space Model Adaptor for Visual RecognitionabstractRecent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior performance due to three main constraints existing in the sequential model: 1) Casual computing is incapable of accessing global context; 2) Long-range forgetting when computing the current hidden states; 3) Weak spatial structural modeling due to the transformed sequential input. To address these issues, we investigate a simple yet powerful vision task adapter for Mamba models, which consists of two functional modules: Adaptor-T and Adapator-S. When solving the hidden states for SSM, we apply a lightweight prediction module Adaptor-T to select a set of learnable locations as memory augmentations to ease long-range forgetting issues. Moreover, we leverage Adapator-S, composed of multi-scale dilated convolutional kernels, to enhance the spatial modeling and introduce the image inductive bias into the feature output. Both modules can enlarge the context modeling in casual computing, as the output is enhanced by the inaccessible features. We explore three usages of Mamba-Adaptor: A general visual backbone for various vision tasks; A booster module to raise the performance of pretrained backbones; A highly efficient fine-tuning module that adapts the base model for transfer learning tasks. Extensive experiments verify the effectiveness of Mamba-Adapter in three settings. Notably, our Mamba-Adapter achieves state-of-the-art performance on the ImageNet and COCO benchmarks. Jiahao Nie 0001, Yujin Tang, Hongshen Zhao |
CVPR | 5 |
| 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationabstractRecent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses object- and semantic-level information to improve the recognition of absent targets. Extensive experiments on gRefCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO, and RefCOCO (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG. Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng 0001, Wankou Yang |
ICCV | 5 |