VLDB 2026 Research / reviewers in the wild / expert
Gongpeng Zhao
dblp:342/7293
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0004-1019-1477ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Face, body and person analysis · 40% Vision and language · 24% Autonomous driving · 11% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Autonomous driving › scenario generation
driving scene generation |
0.9 | 1 | 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic Fusion · ICRA 2025 |
Computer vision › Face, body and person analysis › facial expression analysis
micro-expression analysis |
0.9 | 1 | 2025 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question Answering · ACM Multimedia 2025 |
Computer vision › Vision and language
visual question answering |
0.9 | 1 | 2025 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question Answering · ACM Multimedia 2025 |
Computer vision › Face, body and person analysis
facial expression analysis |
0.8 | 1 | 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration · ACM Multimedia 2024 |
Computer vision › Face, body and person analysis › facial expression analysis › facial expression recognition
micro-expression recognition |
0.8 | 1 | 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize · ACM Multimedia 2024 |
Computer vision › Face, body and person analysis › facial expression analysis › micro-expression analysis
micro-expression spotting |
0.8 | 1 | 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration · ACM Multimedia 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.8 | 1 | 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination Corrector · ACM Multimedia 2024 |
Computer vision › Video understanding and tracking › action detection
temporal action localization |
0.8 | 1 | 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize · ACM Multimedia 2024 |
Computer vision › Vision and language › image captioning › description generation
visual spatial description |
0.8 | 1 | 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination Corrector · ACM Multimedia 2024 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.3 | 1 | 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic Fusion · ICRA 2025 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic Fusion · ICRA 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question Answering · ACM Multimedia 2025 |
Image and video processing › motion estimation
optical flow |
0.2 | 1 | 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal large language model · 1.6optical flow · 1.5boundary calibration · 1.5semantic fusion attention · 0.9occupancy ray sampling · 0.9hierarchical reasoning · 0.9foreground-aware masked loss · 0.9multi-scale feature fusion · 0.8adapter · 0.8VideoMAE · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic FusionabstractAccurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information. In this paper, we propose DualDiff, a dual-branch conditional diffusion model designed to enhance multi-view driving scene generation. We introduce Occupancy Ray Sampling (ORS), a semantic-rich 3D representation, alongside numerical driving scene representation, for comprehensive foreground and background control. To improve cross-modal information integration, we propose a Semantic Fusion Attention (SFA) mechanism that aligns and fuses features across modalities. Furthermore, we design a foreground-aware masked (FGM) loss to enhance the generation of tiny objects. DualDiff achieves state-of-the-art performance in FID score, as well as consistently better results in downstream BEV segmentation and 3D object detection tasks. Haoteng Li, Zezhong Qian, Gongpeng Zhao, Jun Yu 0001, Huazheng Zhou, Longjun Liu |
ICRA | 4 |
| 2025 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question AnsweringabstractThe rise of Multimodal Large Language Models (MLLMs) offers new opportunities for Micro-Expression (ME) analysis. This paper introduces Micro-Expression Visual Question Answering (ME-VQA), a novel task reformulating ME annotations (e.g., emotion categories, action units) into QA pairs. To address key challenges-hardware limitations, context inconsistency, and compositional reasoning gaps-we propose a Relationship-Aware Hierarchical VQA Framework. Our approach leverages mined emotion correlations (e.g., coarse-to-fine label dependencies) and employs a two-stage process: 1) Coarse-grained anchoring for broad emotion categories, and 2) Fine-grained reasoning constrained by coarse outputs and statistical rules. We further optimize efficiency via a dual-phase video sampling strategy: during training, keyframes (onset/apex/offset) and random non-expression frames are used; uniform sampling is applied at inference. Experiments demonstrate significant improvements in answer consistency and accuracy. Lingsi Zhu, Yanjun Chi, Jun Yu 0001, Gongpeng Zhao, Yuefeng Zou, Fengzhao Sun, Xilong Lu |
ACM Multimedia | 4 |
| 2025 | Joint Optic Disc and Cup Segmentation Via KNN-Based Transformer and Deformable Aggregation AttentionabstractDeep learning has advanced medical image segmentation, especially for optic disc and cup detection. While convolutional neural networks (CNNs) struggle with long-range dependencies, Transformer-based architectures have emerged to address this limitation.However, complete replacement of CNNs with Transformers may impair local feature extraction. Additionally, reliance on expert-annotated datasets makes supervised learning costly. To overcome these limitations, we present DAK-Former, a novel self-supervised contrastive learning approach for optic disc and cup segmentation. Our approach introduces a new attention mechanism that combines K-nearest neighbors (KNN) with Transformer components, along with a deformable aggregation attention module to improve global feature representation. Additionally, our method uses multiple types of medical imaging data, including MRI, CT scans, and X-rays. Using contrastive learning, we match encoded queries with a dictionary of encoded keys, allowing the network to learn meaningful unsupervised feature representations. We evaluated DAK-Former on two publicly available fundus image datasets and compared it with state-of-the-art methods. Experimental results show that DAK-Former is highly effective and consistently outperforms existing approaches. Jun Yu 0001, Shuoping Yang, Gongpeng Zhao, Lei Wang 0203 |
IEEE Internet Things J. | 3 |
| 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Zhongpeng Cai, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 3 |
| 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 2 |
| 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination CorrectorabstractVisual Spatial Description (VSD) is an emerging image-to-text task which aims at generating descriptions of the spatial relationships between given objects in an image. In this paper, we apply Retrieval-Augmented Generation (RAG) technology in guiding Multimodal Large Language Models (MLLMs) for the task of VSD, complemented by an Adaptive Hallucination Corrector, and further fine-tuning them to bolster semantic understanding and overall model efficacy. We found that our approach demonstrated higher accuracy and fewer hallucination errors in both spatial relationship classification and visual language description tasks within the VSD task, achieving state-of-the-art results. Jun Yu 0001, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 5 |