VLDB 2026 Research / reviewers in the wild / expert
Chenyu Cao
dblp:292/7045
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Overall-Distinctive GCN for Social Relation Recognition on Videos
Yibo Hu 0005, Chenyu Cao, Fangtao Li, Chenghao Yan, Jinsheng Qi, Bin Wu 0001 |
MMM (1) | 2 |
| 2023 | Social Relation Graph Generation on Untrimmed Video
Yibo Hu 0005, Chenghao Yan, Chenyu Cao, Haorui Wang, Bin Wu 0001 |
MMM (2) | 3 |
| 2022 | FHGN: Frame-Level Heterogeneous Graph Networks for Video Question AnsweringabstractVideo Question Answering (VideoQA), aims to answer the given question based on video grounding, reasoning, and multimodal interacting. Most existing studies usually extract video-level visual and linguistic embeddings separately, then perform spatial-temporal or graph-based reasoning in a single modality. These methods lack the understanding of the interactions among different modalities and ignore the rele-vance of different frames, objects, and modalities to the question. This paper proposes Frame-level Heterogeneous Graph Networks (FHGN), to learn the semantic correlations among different modalities in each frame. Specifically, we construct a uniform frame-level heterogeneous graph to perform both inter- and intra-modality reasoning. Then a two-stage attention module is applied to measure the relevance between the question and different periods, space regions, and modalities. We conduct extensive experiments on a large-scale VideoQA dataset, and the experimental results show that our FHGN achieves the state-of-the-art performance. Jinsheng Qi, Fangtao Li, Ting Bai 0004, Chenyu Cao, Yibo Hu 0005, Bin Wu 0001 |
ICME | 4 |
| 2022 | A Multimodal Approach for Multiple-Relation Extraction in Videos
Weiying Hou, Chenyu Cao, Bin Wu 0001 |
Multim. Tools Appl. | 4 |
| 2021 | Relation-aware Hierarchical Attention Framework for Video Question AnsweringabstractVideo Question Answering (VideoQA) is a challenging video understanding task since it requires a deep understanding of both question and video. Previous studies mainly focus on extracting sophisticated visual and language embeddings, fusing them by delicate hand-crafted networks. However, the relevance of different frames, objects, and modalities to the question are varied along with the time, which is ignored in most of existing methods. Lacking understanding of the the dynamic relationships and interactions among objects brings a great challenge to VideoQA task. To address this problem, we propose a novel Relation-aware Hierarchical Attention (RHA) framework to learn both the static and dynamic relations of the objects in videos. In particular, videos and questions are embedded by pre-trained models firstly to obtain the visual and textual features. Then a graph-based relation encoder is utilized to extract the static relationship between visual objects. To capture the dynamic changes of multimodal objects in different video frames, we consider the temporal, spatial, and semantic relations, and fuse the multimodal features by hierarchical attention mechanism to predict the answer. We conduct extensive experiments on a large scale VideoQA dataset, and the experimental results demonstrate that our RHA outperforms the state-of-the-art methods. Fangtao Li, Ting Bai 0004, Chenyu Cao, Chenghao Yan, Bin Wu 0001 |
ICMR | 3 |
| 2021 | Social Relation Analysis from Videos via Multi-entity ReasoningabstractVideos contain rich semantic information. Analyzing social relations in video semantics can help machines interpret the behavior of human beings. However, most of the work related to social relationship recognition is based on still images, while video-based social relationship analysis tasks are less concerned. Here we propose a Multi-entity Relation Reasoning (MRR) framework that can be used for recognizing or predicting social relations in videos. To capture temporal features and contextual cues in videos, and use richer information to represent the person in the video, we track each person's appearance timeline and design a multi-entity representation method to build a social relationship knowledge graph. Then we use graph attention networks to gather information from the entity's neighborhood. Besides, situation information is helpful to identify relationships, we design a situation information extraction module to generate situation embedding from the video clip. Finally, a decoder is adopted to predict relationships between character entities. We evaluate the model on the MovieGraphs dataset and verify the effectiveness of the proposed framework. Chenghao Yan, Fangtao Li, Chenyu Cao, Bin Wu 0001 |
ICMR | 4 |