Qianyue Bao

dblp:300/8244 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-6025-3881ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Privacy-preserving video anomaly detection via federated learning
Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Qianyue Bao, Lingling Li 0002, Puhua Chen
Knowl. Based Syst.8
2026 ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot Captioning
abstract
Zero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD.
Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 RefZVC: Refinable Zero-Shot Video Captioning by Test-Time Reinforcement Polishing
abstract
Recently, the zero-shot image captioning (zero-shot IC) method based on pre-trained visual language models (VLMs) and large language models (LLMs) has made significant progress. However, how to adapt it to the zero-shot video captioning (zero-shot VC) scenario (without video-text paired supervision) has not been well explored. Inspired by various recent test-time strategies (sacrificing additional test time to improve performance), we try to introduce a new paradigm of Test-time Reinforcement Polishing in zero-shot VC scenario. We take temporal dependency modeling as the starting point and propose a novel framework for Refinable Zero-shot VC, called RefZVC. RefZVC can greatly cover the long-term context of the video and continuously polish and refine the generated captions in a reward-feedback manner. We first design an Adaptive Frame Skipping module (AdaSkip) to skip redundant frames and select diverse keyframe sequences. Subsequently, we propose a Multi-granularity Reinforcement Polishing (MRP) mechanism, which iteratively polishes captions by leveraging Gaussian Kernel Cache (GKC) to capture temporal dynamics, store and reuse relevant historical context. In addition, MRP calculates rewards for generated captions at both the sentence-level and entity-level to achieve test-time polishing. With the MRP mechanism, RefZVC achieves superior zero-shot generalization performance, outperforming previous zero-shot VC methods on benchmarks such as MSVD, MSR-VTT, and VATEX.
Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Lingling Li 0002, Xu Liu 0006
IEEE Trans. Image Process.1
2025 Visual-Language Scene-Relation-Aware Zero-Shot Captioner
abstract
Zero-shot image captioning can harness the knowledge of pre-trained visual language models (VLMs) and language models (LMs) to generate captions for target domain images without paired sample training. Existing methods attempt to establish high-quality connections between visual and textual modalities in text-only pre-training tasks. These methods can be divided into two perspectives: sentence-level and entity-level. Although they achieve effective performance on some metrics, they suffer from hallucinations due to biased associations during training. In this paper, we propose a scene-relation-level pre-training task by considering relations as more valuable modal connection bridges. Based on this, we construct a novel Visual-Language Scene Relation Aware Captioner (SRACap), which expands the ability to predict scene relations while generating captions for images. In addition, SRACap possesses excellent cross-domain zero-shot generalization capability, which is driven by a well-designed scene reinforcement switching pipeline. We introduce a scene policy network to dynamically crop salient regions from images and feed them into a language model to generate captions. We integrate multiple expert CLIP models to form a mixture-of-rewards module (MoR) as a reward source, and deeply optimized SRACap through the policy gradient algorithm in the zero-shot inference stage. With the iteration of scene reinforcement switching, SRACap can gradually refine the generated caption details while maintaining high semantic consistency across visual-linguistic modalities. We conduct extensive experiments on multiple standard image captioning benchmarks, showing that SRACap can accurately understand scene structures and generate high-quality text, significantly outperforming other zero-shot inference methods.
Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Knowledge-Driven Compositional Action Recognition
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006
Pattern Recognit.4
2025 Anomaly-Led Prompting Learning Caption Generating Model and Benchmark
abstract
Video anomaly detection (VAD) is an important intelligent system application, but most current research views it as a coarse binary classification task that lacks a fine-grained understanding of abnormal video sequences. We explore a new task for video anomaly analysis called Comprehensive Video Anomaly Caption (CVAC), which aims to generate comprehensive textual captions (containing scene information such as time, location, anomalous subject, anomalous behavior, etc.) for surveillance videos. CVAC is more consistent with human understanding than VAD, but it has not been well explored. We constructed a large-scale benchmark CVACBench to lead this research. For each video clip, we provide 6 fine-grained annotations, including scene information and abnormal keywords. A new evaluation metric Abnormal-F1 (A-F1) is also proposed to more accurately evaluate the caption generation performance of the model. We also designed a method called Anomaly-Led Generating Prompting Transformer (AGPFormer) as a baseline. In AGPFormer, we introduce an anomaly-led language modeling mechanism (Anomaly-Led MLM, AMLM) to focus on anomalous events in videos. To achieve more efficient cross-modal semantic understanding, we design the Interactive Generating Prompting (IGP) module and Scene Alignment Prompting (SAP) module to explore the divide between video and text modalities from multiple perspectives, and to improve the model's performance in understanding and reasoning about the complex semantics of videos. We conducted experiments on CVACBench by using traditional caption metrics and the proposed metrics, and the experimental results demonstrate the effectiveness of AGPFormer in the field of anomaly caption.
Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Baoliang Chen
IEEE Trans. Multim.1
2025 Chain-of-Situation Aware Progressive Inference Learning
abstract
The grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL.
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Multi-Grained Gradual Inference Model for Multimedia Event Extraction
abstract
With the development of multimedia technology, events are usually presented in multimedia forms, thus multimedia event extraction (MEE) has become more and more important. Existing MEE works usually use simple strategies to align two modalities, making it difficult to precisely extract events and arguments in complex multimedia documents. To address this problem, we propose a novel Multi-grained Gradual Inference Model (MGIM) that focuses on inferring and interpreting events in complex multimedia structures in a coarse-to-fine manner. To efficiently integrate textual and visual modalities, we design a Coarse-grained Alignment (CA) module, which represents the two modalities in a graph structure and performs coarse-grained alignment. Based on the CA module, we further propose a Fine-grained Inference module (FI) that fine-grained aligns text and image by performing multiple rounds of gradual inference. MGIM provides a comprehensive interpretation of multimedia events at two information granularities (coarse and fine). Extensive experiments on the M2E2 dataset demonstrate the effectiveness of MGIM.
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006
IEEE Trans. Circuits Syst. Video Technol.4
2024 A Knowledge-Based Hierarchical Causal Inference Network for Video Action Recognition
abstract
Currently, existing action recognition methods mainly use a data-driven method to extract spatio-temporal representations of actions for recognition. However, this method may face performance bottlenecks. At the same time, existing action recognition methods are easily affected by the bias of scene information and object information in videos. In order to explore the essential causal relationship between factors and remove bias in action recognition, we introduce the theory of causal inference into the field of action recognition and propose a Knowledge-based Hierarchical Causal Inference Network (KHCIN) to help us step toward a new direction of inference in action recognition. First, we construct a Knowledge-based Hierarchical Causal Graph (KHCG) to structurally represent the scene, object and motion knowledge of a video. Then, in the model inference stage, we perform factual causal inference on a video on the constructed KHCG, and then deploy counterfactual inference on the Direct Content Hierarchy (DCH) and Indirect Interaction Hierarchy (IIH) in the KHCG. For DCH, we intervene in the model at the decision level to highlight bias errors in the model predictions. For the IIH, we focus on intervening in the feature modelling process. The biased interactions are revealed by interrupting the information communication in the feature space. By comparing the results of factual and counterfactual inference, we can easily expose the biased information in the original representations and eliminate them. Driven by counterfactual causal inference, our approach can significantly improve the performance of action recognition while improving model explainability. Extensive experiments demonstrate the effectiveness of this method. We hope that KHCIN can provide some new ideas for better introduction of causal inference theory in the action recognition community in the future.
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Lingling Li 0002, Yuwei Guo 0001, Puhua Chen
IEEE Trans. Multim.4
2022 In-Place Gestures Classification via Long-term Memory Augmented Network
abstract
In-place gesture-based virtual locomotion techniques enable users to control their viewpoint and intuitively move in the 3D virtual environment. A key research problem is to accurately and quickly recognize in-place gestures, since they can trigger specific movements of virtual viewpoints and enhance user experience. However, to achieve real-time experience, only short-term sensor sequence data (up to about 300ms, 6 to 10 frames) can be taken as input, which actually affects the classification performance due to limited spatiotemporal information. In this paper, we propose a novel long-term memory augmented network for in-place gestures classification. It takes as input both short-term gesture sequence samples and their corresponding long-term sequence samples that provide extra relevant spatio-temporal information in the training phase. We store long-term sequence features with an external memory queue. In addition, we design a memory augmented loss to help cluster features of the same class and push apart features from different classes, thus enabling our memory queue to memorize more relevant long-term sequence features. In the inference phase, we input only short-term sequence samples to recall the stored features accordingly, and fuse them together to predict the gesture class. We create a large-scale in-place gestures dataset from 25 participants with 11 gestures. Our method achieves a promising accuracy of 95.1% with a latency of 192ms, and an accuracy of 97.3% with a latency of 312ms, and is demonstrated to be superior to recent in-place gesture classification techniques. User study also validates our approach. Our source code and dataset will be made available to the community.
Lizhi Zhao, Xuequan Lu, Qianyue Bao, Meili Wang 0001
ISMAR3
2022 Hierarchical Scene Normality-Binding Modeling for Anomaly Detection in Surveillance Videos
abstract
Anomaly detection in surveillance videos is an important topic in the multimedia community, which requires efficient scene context extraction and the capture of temporal information as a basis for decision. From the perspective of hierarchical modeling, we parse the surveillance scene from global to local and propose a Hierarchical Scene Normality-Binding Modeling framework (HSNBM) to handle anomaly detection. For the static background hierarchy, we design a Region Clustering-driven Multi-task Memory Autoencoder (RCM-MemAE), which can simultaneously perform region segmentation and scene reconstruction. The normal prototypes of each local region are stored, and the frame reconstruction error is subsequently amplified by global memory augmentation. For the dynamic foreground object hierarchy, we employ a Scene-Object Binding Frame Prediction module (SOB-FP) to bind all foreground objects in the frame with the prototypes stored in the background hierarchy according their positions, thus fully exploit the normality relationship between foreground and background. The bound features are then fed into the decoder to predict the future movement of the objects. With the binding mechanism between foreground and background, HSNBM effectively integrates the "reconstruction" and "prediction" tasks and builds a semantic bridge between the two hierarchies. Finally, HSNBM fuses the anomaly scores of the two hierarchies to make a comprehensive decision. Extensive empirical studies on three standard video anomaly detection datasets demonstrate the effectiveness of the proposed HSNBM framework.
Qianyue Bao, Fang Liu 0001, Yang Liu 0349, Licheng Jiao, Xu Liu 0006, Lingling Li 0002
ACM Multimedia1
2021 MRTA: Multi-Resolution Training Algorithm for Multitemporal Semantic Change Detection
abstract
The multitemporal semantic change detection challenge track (Track MSD) in the 2021 Data Fusion Contest is to extract the land cover changes of the US state of Maryland from 2013 to 2017, but only low-resolution label data is provided. We present a multi-resolution training algorithm (MRTA) to alleviate the overfitting of the model on the coarse labels. First, using low-resolution coarse labels to train FCN, the average IoU can reach 0.5253. Generating pseudo-labels using this network, they are combined with coarse labels to form a multi-resolution label combination, and perform iterative fine-tuning and retrain. After that, by analyzing the loss and gain indicators of specific categories, it was found that the discrimination effect of the water area was poor, so a strong classifier was trained for the water area. We also implemented strategies such as model voting and weighted training to improve model performance. Finally, our method achieves 0.6445 mIoU on the test set, ranking 3rd in Track 2 of the IEEE Data Fusion Contest of 2021.
Qianyue Bao, Yang Liu 0349, Zixiao Zhang, Dafan Chen, Yuting Yang 0008, Licheng Jiao, Fang Liu 0001
IGARSS1