VLDB 2026 Research / reviewers in the wild / expert
Yiming Li 0008
dblp:181/2877-8
· DBLP profile ↗
6ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-1958-8698ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatial-Temporal Prior Knowledge Guidance for Long-term Action AnticipationabstractFor long-term action anticipation (LAA), the primary focus is on understanding the observed video content and anticipating future actions, including both the names of upcoming actions and their corresponding durations. This requires the model to fully grasp the patterns of action transitions. To facilitate this, we employ spatial-temporal prior knowledge to guide the LAA model in capturing these transition patterns, which is referred to as explicit learning. Additionally, we use a transformer structure which incorporates a parallel decoding mechanism in which mitigates error accumulation and can be regarded as implicit learning. Consequently, we propose a novel model that integrates explicit and implicit learning approaches, combining the advantages of both. On the benchmarks for long-term action anticipation, our method achieves state-of-the-art results on the 50Salads and Breakfast. Yiming Li 0008, Miao Ji, Sisi You, Bing-Kun Bao |
ICME | 1 |
| 2025 | Graph Prompts: Adapting Video Graph for Video Question AnsweringabstractDue to the dynamic nature in videos, it is evident that perceiving and reasoning about temporal information are the key focus of Video Question Answering (VideoQA). In recent years, several methods have explored relationship-level temporal modeling with graph-structured video representation. Unfortunately, these methods heavily rely on the question text, thus making it challenging to perceive and reason about video content that is not explicitly mentioned in the question. To address the above challenge, we propose Graph Prompts-based VideoQA (GP-VQA), which adopts a video-based graph structure for enhanced video understanding. The proposed GP-VQA contains two stages, i.e., pre-training and prompt tuning. In pre-training, we define the pretext task that requires GP-VQA to reason about the randomly masked nodes or edges in the video graph, thus prompting GP-VQA to learn the reasoning ability with video-guided information. In prompt-tuning, we organize the textual question into question graph and implement message passing from video graph to question graph, therefore inheriting the video-based reasoning ability from video graph completion to VideoQA. Extensive experiments on various datasets have demonstrated the promising performance of GP-VQA. Yiming Li 0008, Xiaoshan Yang, Bing-Kun Bao, Changsheng Xu |
IJCAI | 1 |
| 2023 | Iterative Learning with Extra and Inner Knowledge for Long-tail Dynamic Scene Graph GenerationabstractDynamic scene graphs have become a powerful tool for higher-level visual understanding tasks, and the interest in dynamic scene graph generation (dynamic SGG) is grown over time. Recently, numbers of existing methods achieve significant progress in dynamic SGG by capturing temporal information with transformer or recurrent network structures. However, most existing methods only focus on predicting the head predicates, which ignore the long-tail phenomenon, thus the tail predicates are hard to be recognized. In this paper, we propose a novel method named Iterative Learning with Extra and Inner Knowledge (I2LEK) to address the long-tail problem in dynamic SGG. The extra knowledge is obtained from commonsense, while inner knowledge is defined as the temporal evolution patterns of visual relationships. Specifically, we introduce extra knowledge to enrich the representations of predicates in the spatial dimension and adopt inner knowledge to implement knowledge sharing in the temporal dimension. With enriched representations and shared knowledge, I2LEK can accurately predict both the tail and head predicates. Moreover, an iterative learning strategy is proposed to fuse the extra knowledge, inner knowledge, and spatial-temporal context contained in videos, which further enhances the model's understanding of visual relationships. Our experimental results on the public Action Genome dataset demonstrate that our model achieves state-of-the-art performance. Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
ACM Multimedia | 1 |
| 2023 | Zero-Shot Predicate Prediction for Scene Graph ParsingabstractThe scene graph is a structured semantic representation of an image, which represents objects and relationships with vertices and edges, respectively. Since it is impossible to manually label all potential relationships in the real world, some previous methods try to apply the zero-shot method for scene graph generation. However, existing methods take triplet (i.e., hsubject-predicate-objecti) as the basic unit of a relationship. Each element (i.e., subject, predicate, or object) of the unseen relationship is actually seen in the training data. Therefore, they ignore the unseen predicate. To predict the unseen predicate, we introduce a novel task named zero-shot predicate prediction, which is crucial to extending existing scene graph generation methods to recognize more relationship classes. The new task is challenging and cannot be simply resolved through conventional zero-shot learning methods because there is a large intra-class variation of each predicate. Firstly, the large intra-class variation leads to the difficulty of computing the discriminative instancelevel feature of the predicate class. Secondly, the large intraclass variation also brings more difficulties when knowledge is transferred from seen classes to unseen classes. For the first challenge, we propose distilling lexical knowledge of different objects and construct multi-modal representations of pairwise objects to reduce the intra-class variation of the predicate. To respond to the second challenge, we build a compact semantic space where the representations of unseen classes are reconstructed based on the seen classes for zero-shot predicate classification. We evaluate the proposed method on the public dataset Visual Genome. The extensive experiment results under the zeroshot/few-shot/supervised settings demonstrate the effectiveness of the proposed method. Yiming Li 0008, Xiaoshan Yang, Xuhui Huang, Zhe Ma 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2022 | Dynamic Scene Graph Generation via Anticipatory Pre-trainingabstractHumans can not only see the collection of objects in visual scenes, but also identify the relationship between objects. The visual relationship in the scene can be abstracted into the semantic representation of a triple (subject, predicate, object) and thus results in a scene graph, which can convey a lot of information for visual understanding. Due to the motion of objects, the visual relationship between two objects in videos may vary, which makes the task of dynamically generating scene graphs from videos more complicated and challenging than the conventional image-based static scene graph generation. Inspired by the ability of humans to infer the visual relationship, we propose a novel anticipatory pre-training paradigm based on Transformer to explicitly model the temporal correlation of visual relationships in different frames to improve dynamic scene graph generation. In pre-training stage, the model predicts the visual relationships of current frame based on the previous frames by extracting intra-frame spatial information with a spatial encoder and inter-frame temporal correlations with a progressive temporal encoder. In the fine-tuning stage, we reuse the spatial encoder and the progressive temporal encoder while the information of the current frame is combined for predicting the visual relationship. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Action Genome dataset. Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
CVPR | 1 |
| 2020 | Structured Neural Motifs: Scene Graph Parsing via Enhanced Context
Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
MMM (2) | 1 |