Kaifeng Gao

dblp:247/1121 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 32% Video understanding and tracking · 31% Trustworthy machine learning · 11%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.722025
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing · ICML 2025
Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards · CVPR 2025
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
video scene graph generation
1.222023
Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation · ACM Multimedia 2023
Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs · CVPR 2022
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
video visual relation detection
1.222023
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection · ICLR 2023
Video Relation Detection via Tracklet based Visual Transformer · ACM Multimedia 2021
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
autoregressive generation
0.912025
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data · ICML 2025
Machine learning › Generative modeling
score-based model
0.912025
Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data · ICML 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing · ICML 2025
Machine learning › Generative modeling
video generation
0.912025
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing · ICML 2025
Machine learning › Trustworthy machine learning
fairness
0.712023
Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation · ACM Multimedia 2023
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
open-vocabulary video visual relationship detection
0.712023
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection · ICLR 2023
Computer vision › Vision and language
cross-modal alignment
0.612022
Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives · EMNLP 2022
Computer vision › Video understanding and tracking
video question answering
0.612022
Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives · EMNLP 2022
Computer vision › Video understanding and tracking
object tracking
0.512021
Video Relation Detection via Tracklet based Visual Transformer · ACM Multimedia 2021
Machine learning › Graph learning › heterogeneous graph
bipartite graph
0.212022
Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs · CVPR 2022

Methods — techniques the papers use, named apart from their topics

score-based model · 0.9policy optimization · 0.9density estimation · 0.9causal generation · 0.9cache-sharing · 0.9branch-based sampling · 0.9backward progressive training · 0.9autoregressive generation · 0.9prompt tuning · 0.7motion cues · 0.7
YearPublicationVenuePosition
2026 Generalized Visual Relation Detection With Diffusion Models
abstract
Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance, they are all restricted to pre-defined relation categories, while failing to consider thesemantic ambiguitycharacteristic of visual relations. Unlike objects, the appearance of visual relations is always subtle and can be described by multiple predicate words from different perspectives, e.g., “ride” can be depicted as “race” and “sit on”, from the sports and spatial position views, respectively. To this end, we propose to model visual relations as continuous embeddings, and design diffusion models to achieve generalized VRD in a conditional generative manner, termed Diff-VRD. We model the diffusion process in a latent space and generate all possible relations in the image as an embedding sequence. During the generation, the visual and text embeddings of subject-object pairs serve as conditional signals and are injected via cross-attention. After the generation, we design a subsequent matching stage to assign the relation words to subject-object pairs by considering their semantic similarities. Benefiting from the diffusion-based generative process, our Diff-VRD is able to generate visual relations beyond the pre-defined category labels of datasets. To properly evaluate this generalized VRD task, we introduce two evaluation metrics, i.e., text-to-image retrieval and SPICE PR Curve inspired by image captioning. Extensive experiments in both human-object interaction (HOI) detection and scene graph generation (SGG) benchmarks attest to the superiority and effectiveness of Diff-VRD.
Kaifeng Gao, Hanwang Zhang, Jun Xiao 0001, Yueting Zhuang, Qianru Sun
IEEE Trans. Circuits Syst. Video Technol.1
2025 Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
abstract
Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1.
Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001
CVPR6
2025 Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
abstract
With the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an autoregressive manner, i.e., generating subsequent clips conditioned on the last frame(s) of the previous clip. However, existing autoregressive VDMs are highly inefficient and redundant: The model must re-compute all the conditional frames that are overlapped between adjacent clips. This issue is exacerbated when the conditional frames are extended autoregressively to provide the model with long-term context. In such cases, the computational demands increase significantly (i.e., with a quadratic complexity w.r.t. the autoregression step). In this paper, we propose **Ca2-VDM**, an efficient autoregressive VDM with **Ca**usal generation and **Ca**che sharing. For **causal generation**, it introduces unidirectional feature computation, which ensures that the cache of conditional frames can be precomputed in previous autoregression steps and reused in every subsequent step, eliminating redundant computations. For **cache sharing**, it shares the cache across all denoising steps to avoid the huge cache storage cost. Extensive experiments demonstrated that our Ca2-VDM achieves state-of-the-art quantitative and qualitative video generation results and significantly improves the generation speed. Code is available: https://github.com/Dawn-LX/CausalCache-VDM
Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016
ICML1
2025 Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data
abstract
Machine learning models often perform well on tabular data by optimizing average prediction accuracy. However, they may underperform on specific subsets due to inherent biases and spurious correlations in the training data, such as associations with non-causal features like demographic information. These biases lead to critical robustness issues as models may inherit or amplify them, resulting in poor performance where such misleading correlations do not hold. Existing mitigation methods have significant limitations: some require prior group labels, which are often unavailable, while others focus solely on the conditional distribution $P(Y|X)$, upweighting misclassified samples without effectively balancing the overall data distribution $P(X)$. To address these shortcomings, we propose a latent score-based reweighting framework. It leverages score-based models to capture the joint data distribution $P(X, Y)$ without relying on additional prior information. By estimating sample density through the similarity of score vectors with neighboring data points, our method identifies underrepresented regions and upweights samples accordingly. This approach directly tackles inherent data imbalances, enhancing robustness by ensuring a more uniform dataset representation. Experiments on various tabular datasets under distribution shifts demonstrate that our method effectively improves performance on imbalanced data.
Yunze Tong, Fengda Zhang, Kaifeng Gao, Pengfei Lyu, Jun Xiao 0001, Kun Kuang 0001
ICML4
2023 Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Kaifeng Gao, Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Qianru Sun
ICLR1
2023 Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation
abstract
Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to the inherently biased distribution and missing annotations in the training data, current VidSGG methods have been found to perform poorly on less-represented predicates. In this paper, we propose an explicit solution to address this under-explored issue by supplementing missing predicates that should be included in the ground-truth annotations. Dubbed Trico, our method seeks to supplement the missing predicates that are supposed to appear in the ground-truth annotations, by exploring three complementary spatio-temporal correlations. Guided by these correlations, the missing labels can be effectively supplemented thus achieving an unbiased predicate predictions. We validate the effectiveness of Trico on the most widely used VidSGG datasets, i.e., VidVRD and VidOR. Extensive experiments demonstrate the state-of-the-art performance achieved by Trico, particularly on those tail predicates. The code is available in the supplementary material.
Kaifeng Gao, Yawei Luo, Tao Jiang 0042, Fei Gao 0014, Jian Shao 0001, Jun Xiao 0001
ACM Multimedia2
2022 Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs
abstract
Today's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The ground-truth predicate labels for proposals are partially correct. 2) They break the high-order relations among different predicate instances of a same subject-object pair. 3) VidSGG performance is upper-bounded by the quality of the proposals. To this end, we propose a new classification-then-grounding framework for VidSGG, which can avoid all the three overlooked drawbacks. Meanwhile, under this framework, we reformulate the video scene graphs as temporal bipartite graphs, where the entities and predicates are two types of nodes with time slots, and the edges denote different semantic roles between these nodes. This formulation takes full advantage of our new framework. Accordingly, we further propose a novel BIpartite Graph based SGG model: BIG. It consists of a classification stage and a grounding stage, where the former aims to classify the categories of all the nodes and the edges, and the latter tries to localize the temporal location of each relation instance. Extensive ablations on two VidSGG datasets have attested to the effectiveness of our framework and BIG. Code is available at https://github.com/Dawn-LX/VidSGG-BIG.
Kaifeng Gao, Long Chen 0016, Yulei Niu, Jian Shao 0001, Jun Xiao 0001
CVPR1
2022 Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives
abstract
Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at different levels of abstraction.Existing efforts mainly focus on designing sophisticated architectures while utilizing frame-or object-level visual representations.In this paper, we reconsider the multi-modal alignment in VideoQA from feature and sample perspectives to achieve better performance.From the view of feature, we break down the video into trajectories and first leverage trajectory feature in VideoQA to enhance the alignment between two modalities.Moreover, we adopt a heterogeneous graph architecture and design a hierarchical framework to align both trajectory-level and frame-level visual feature with language feature.In addition, we found that VideoQA models are largely dependent on language priors and always neglect visuallanguage interactions.Thus, two effective yet portable training augmentation strategies are designed to strengthen the cross-modal correspondence ability of our model from the view of sample.Extensive results show that our method outperforms all state-of-the-art models on the challenging NExT-QA benchmark.* Long Chen is the corresponding author.… objects trajectories Q: Why did the boy in orange hold a ball on his head?A: want to throw the ball.
Shaoning Xiao, Long Chen 0016, Kaifeng Gao, Yi Yang 0001, Jun Xiao 0001
EMNLP3
2021 Video Relation Detection via Tracklet based Visual Transformer
abstract
Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA[7] and deepSORT [27] to generate tracklet proposals. Then we perform VidVRD in a tracklet-based manner without any pre-cutting operations. Specifically, we design a tracklet-based visual Transformer. It contains a temporal-aware decoder which performs feature interactions between the tracklets and learnable predicate query embeddings, and finally predicts the relations. Experimental results strongly demonstrate the superiority of our method, which outperforms other methods by a large margin on the Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021. Codes are released at https://github.com/Dawn-LX/VidVRD-tracklets.
Kaifeng Gao, Long Chen 0016, Jun Xiao 0001
ACM Multimedia1
2020 ARBF: adaptive radial basis function interpolation algorithm for irregularly scattered point sets
Kaifeng Gao, Gang Mei, Salvatore Cuomo, Francesco Piccialli, Nengxiong Xu
Soft Comput.1