VLDB 2026 Research / reviewers in the wild / expert
Dongpan Chen
dblp:331/0266
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-2011-100XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASK-HOI: Affordance-Scene Knowledge Prompting for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection task aims to learn how humans interact with surrounding objects by inferring fine-grained triples of$\left\langle \rm {\emph {human, action, object}} \right\rangle$, which plays a vital role in computer vision tasks such as human-centered scene understanding and visual question answering. However, HOI detection suffers from class long-tailed distributions and zero-shot problems. Current methods typically identify HOI only from input images or label spaces in a data-driven manner, lacking sufficient knowledge prompts, and consequently limits their potential for real-world scenes. Hence, to fill this gap, this paper introduces affordance and scene knowledge as prompts on different granularities to the HOI detector to improve its recognition ability. Concretely, we first construct a large-scale affordance-scene knowledge graph, named ASKG, whose knowledge can be divided into two categories according to the fields of image information, i.e., the knowledge related to affordances of object instances and the knowledge associated with the scene. Subsequently, the knowledge of affordance and scene specific to the input image is extracted by an ASKG-based prior knowledge embedding module. Since this knowledge corresponds to the image at different granularities, we then propose an instance field adaptive fusion module and a scene field adaptive fusion module to enable visual features fully absorb the knowledge prompts. These two encoded features of different fields and knowledge embeddings are finally fed into a proposed HOI recognition module to predict more accurate HOI results. Extensive experiments on both HICO-DET and V-COCO benchmarks demonstrate that the proposed method leads to competitive results compared with the state-of-the-art methods. Dongpan Chen, Dehui Kong, Junna Gao, Qianxing Li |
IEEE Trans. Multim. | 1 |
| 2025 | MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsabstractAffordance refers to the interactable functional properties of an object, and affordance segmentation aims to pixel-level segment the object functional parts in a given image, which is crucial for various interactive vision tasks. Existing methods address the affordance segmentation problem by utilizing only image features, they can hardly solve the problems of interference between adjacent object pixels in complex scenes, and inability to generalize to the open-world. To tackle these problems, we propose a novel open-vocabulary affordance segmentation task and a benchmark dataset, and propose an approach with object shape mask prompts. The mask is used as prior for different granularity visual feature enhancement and fine-grained text prompt embedding. Specifically, we first propose a mask prompt generation module, which generates refined object shape masks, as well as text prompts for mask-focused regions. Based on the masks, we propose a mask prompt feature enhancement module. It uses masks to encode instance features, and then aggregates them with global features to enhance the visual feature representation. The enhanced visual features are combined with text prompts of different granularity to generate class-agnostic affordance mask proposals. We finally classify these proposals in a proposed affordance prediction module. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior performance on a proposed benchmark dataset. Our approach is also competitive on other open-vocabulary part segmentation datasets. Dongpan Chen, Dehui Kong |
AAAI | 1 |
| 2025 | 3d human pose estimation based on conditional dual-branch diffusion
Zhuowei Bai, Dehui Kong, Dongpan Chen, Qianxing Li |
Multim. Syst. | 4 |
| 2025 | Multi-Anchor Offset Representation Based Coarse-to-Fine Diffusion Model for Human Pose Estimationabstract3D human pose estimation (3DHPE) in images aims at estimating 3D joint positions from images. The existing 3DHPE methods usually define the loss function as the error measured by Euclidean distance between the locations of the predicted joints and the ground truth of joints, which confuses two different kinds of errors with obviously different characteristics and should not be processed equally: the error caused by different pose structures and the others. However, The existing human pose representations are not suitable to distinguish these two kinds of errors. In order to tackle this problem, we propose a novel Multi-Anchor Offset Representation (MAOR) for human pose, which locates the position of each joint using its offsets from a group of selected high-precision joints named Multi-Anchor. Making use of MAOR, the pose error related to the distortion of spatial structure can be measured independently from other errors, which is helpful in promoting the accuracy of pose estimation. We then propose a novel MAOR-based coarse-to-fine diffusion model (MAOR-DiffPose) for pose estimation, which optimizes different types of errors of poses step by step. Firstly, a MAOR-based Denoising Process (MDP) is devised to explicitly optimize spatial structures of 3D poses by using MAOR to describe poses and improves the inductive learning ability of MAOR-DiffPose by extracting view-independent features. Secondly, a Joint Coordinate Denoising Process assisted by MAOR (JCDPaM) is devised to expand the input features meaningfully by combining MAOR with the pose representation based on joint coordinates and optimize the joint coordinates of 3D poses with the assistance of MAOR. MAOR-DiffPose realizes accurate 3DHPE by iterating MDP and JCDPaM modules. Comprehensive experimental results on widely used 3DHPE benchmarks Human3.6M and MPI-INF-3DHP show that the proposed method achieves competitive performance compared with the state-of-the-art methods. Qianxing Li, Dehui Kong, Dongpan Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Dual-Branch Knowledge Enhancement Network with Vision-Language Model for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, pre-trained Vision-Language Models (VLM) have shown their great recognition ability in HOI detection task. However, these VLM based methods are struggle to transfer knowledge to achieve desired performance. To this end, we propose a Dual-Branch Knowledge Enhancement Network with VLM (DBKEN-VLM) within the two-stage paradigm to enhance the effectiveness of VLM. Specifically, we propose a semantic mining decoder to supplement contextual and action-related semantic information into our model. It forms a dual-branch knowledge enhancement network with spatial guided decoder. Furthermore, we propose a two-level fusion strategy for the dualbranch network to facilitate better knowledge transfer of VLM. One is feature-level fusion, producing more instructive interaction features; another is decision-level fusion, further enhancing the capability of VLM for HOI detection. The proposed method achieves competitive performance compared to recent methods on two benchmark datasets, HICO-DET and V-COCO. Guangpu Zhou, Dehui Kong, Dongpan Chen, Zhuowei Bai |
IJCNN | 4 |
| 2024 | ADOSMNet: a novel visual affordance detection network with object shape mask guided feature encoders
Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
Multim. Tools Appl. | 1 |
| 2024 | OASNet: Object Affordance State Recognition Network With Joint Visual Features and Relational Semantic EmbeddingsabstractTraditional affordance learning tasks aim to understand object’s interactive functions in an image, such as affordance recognition and affordance detection. However, these tasks cannot determine whether the object is currently interacting, which is crucial for many follow-up tasks, including robotic manipulation and planning task. To fill this gap, this paper proposes a novel object affrodance state (OAS) recognition task, i.e., simultaneously recognizing an object’s affordances and the partner objects that are interacting with it. Accordingly, to facilitate the application of deep learning technology, an OAS recognition task related dataset OAS10k is constructed by collecting and labeling over 10k images. In the dataset, a sample is defined as a set of an image and its OAS labels, each label is represented as$\left \langle{ \rm {\textit {subject, subject's affrodance, interacted object}} }\right \rangle $. These triplet labels have rich relational semantic information, which can improve OAS recognition performance. We hence construct a directed OAS knowledge graph of affordance states, and extract an OAS matrix from it for modelling the semantic relationships of the triplets. Based on the matrix, we propose an OAS recognition network (OASNet), which utilizes GCN to capture the relational semantic embeddings, and uses a transformer to fuse them with the visual features from an image to recognize the affordance states of objects in the image. Experimental results on OAS10k dataset and other triplet label recognition datasets demonstrate that the proposed OASNet achieves the best performance compared to the state-of-the-art methods. The dataset and codes will be released onhttps://github.com/mxmdpc/OAS. Dongpan Chen, Dehui Kong, Lichun Wang 0002, Junna Gao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | A Survey of Visual Affordance Recognition Based on Deep LearningabstractVisual affordance recognition is an important research topic in robotics, human-computer interaction, and other computer vision tasks. In recent years, deep learning-based affordance recognition methods have achieved remarkable performance. However, there is no unified and intensive survey of these methods up to now. Therefore, this article reviews and investigates existing deep learning-based affordance recognition methods from a comprehensive perspective, hoping to pursue greater acceleration in this research domain. Specifically, this article first classifies affordance recognition into five tasks, delves into the methodologies of each task, and explores their rationales and essential relations. Second, several representative affordance recognition datasets are investigated carefully. Third, based on these datasets, this article provides a comprehensive performance comparison and analysis of the current affordance recognition methods, reporting the results of different methods on the same datasets and the results of each method on different datasets. Finally, this article summarizes the progress of affordance recognition, outlines the existing difficulties and provides corresponding solutions, and discusses its future application trends. Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Big Data | 1 |
| 2022 | Deep Sparse Representation Based Image Restoration With Denoising PriorabstractAs a powerful statistical signal modeling technique, sparse representation has been widely used in various image restoration (IR) applications. The sparsity-based methods have achieved leading performance in the past few decades. However, in recent years it has been surpassed by other methods, especially the recent deep learning based methods. In this paper, we address the question that whether sparse representation can be competitive again. The way we answer this question is to redesign it with a deep architecture. To be specific, we propose an end-to-end deep architecture that follows the process of the sparse representation based IR. In particular, we learn a sparse convolutional dictionary to replace the traditional dictionary, and a convolutional neural network (CNN) denoising prior to replace the image prior. Through end-to-end training, the parameters in convolutional dictionary and CNN denoiser can be jointly optimized. Experimental results on several representative IR tasks, including image denoising, deblurring and super-resolution, demonstrate that the proposed deep network can achieve superior performance against state-of-the-art model-based and learning-based methods. Wei Xu 0059, Qing Zhu 0004, Na Qi, Dongpan Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |