Sisi You

dblp:256/2432 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2026
0009-0004-2173-1694ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Coffee-Mate: Assisting keyframe selection via self-distillation for video question answering
Sisi You
Pattern Recognit.2
2026 Question Understanding and Temporality Guiding for Video Question Answering
abstract
Video Question Answering (VideoQA) aims to answer a question based on the content of a given video. Recent methods adapt image-text pre-trained models to the VideoQA task by designing learnable temporal modules within the image encoder. However, these methods struggle to fully comprehend the questions and effectively extract temporal information due to 1) over-reliance on candidate answers and 2) lack of explicit temporal modeling. Specifically, since the question is fixed in different question-answer pairs, existing models tend to focus on the varying candidate answers. Moreover, existing methods merely utilize the classification loss to constrain the confidence of candidate answers, failing to differentiate the effectiveness of temporal information and to explicitly guide temporal modeling. In this paper, we introduce the Question Understanding and Temporality Guiding (QU-TG) method to address the aforementioned limitations. To reduce over-reliance on candidate answers, we propose providing diverse questions through question selection and enhancing the model's comprehensive understanding of questions through question-video matching. To conduct explicit temporal modeling guiding, we propose negative video prevention and positive video guidance to conduct explicit temporal modeling guiding. Negative video prevention incorporates a prevention loss to discourage the model from making predictions based on erroneous temporal cues, whereas positive video guidance utilizes classification loss to encourage the model to derive correct answers from positive videos. Extensive experiments on the NExT-QA, IntentQA, STAR-QA, and Causal-VidQA datasets demonstrate the effectiveness and generalization of our method.
Sisi You, Bing-Kun Bao
IEEE Trans. Multim.2
2026 FoodDiff: A Collaborative Relationship Perception Framework for Food Image Synthesis Using Diffusion Models
abstract
Food image generation is a typical application of text-to-image (T2I) models. The core difference between food image synthesis and other T2I tasks is that there exist complex collaborative relationships among ingredients, cooking actions, and food images, which determine the appearance of dishes. However, existing food image generation models generally ignore or fail to sufficiently utilize such collaborative relationships, which hinders the model from precisely perceiving the shapes and details of food. Furthermore, the pre-training distribution of T2I models is usually noisy and differs from the user-preferred food feature distributions, resulting in deviations from human aesthetics. To address the above issues, we proposeFoodDiff, a collaborative relationship-aware diffusion model for food image generation, which consists of three key components: (1) To perceive collaborative relationships, we propose a collaborative relation module to extract these relations and inject them into the image generation process. (2) To sufficiently interact with the relationships between recipe semantics and food representations, we propose a recipe fusion fine-tuning module to precisely fuse recipe semantics with visual features and fine-tune the pre-trained model. (3) To make the pre-training feature distribution conform to human preference, we introduce an image reward feedback mechanism to optimize the aesthetics of food images. In addition, we propose a high-quality food dataset named Food-Aesthetic with exquisite plates and elaborate annotations. Extensive experiments and human evaluations show that FoodDiff has superior image aesthetics and semantic consistency.
Mengling Xu, Sisi You, Bing-Kun Bao
IEEE Trans. Multim.2
2025 InstantPainting: Expanding GANs for Efficient Text-Conditioned Image Generation Platform
abstract
Text-conditioned image generation enables cross-modal comprehension. Recent emergence of many platforms have found applications in diverse domains like assisted designing and video gaming. However, there still exist challenges in existing platforms due to their expensive training and time-consuming generation processes. In this paper, we introduce an efficient text-conditioned image generation platform, termed InstantPainting. Unlike existing platforms based on large-scale pre-trained diffusion models, InstantPainting expands generative adversarial networks (GANs) to achieve efficient generation by using only about three percent pre-training data of other platforms. Compared to existing platforms, InstantPainting achieves the following functions at a very low deployment cost and approximately 4 to 5 times faster generation speeds: (1) Multi-category and multi-size image generation (2) Image stylization and controlled generation (3) Creative generation, including the generation of poetry pictures and counterfactual images. The proposed platform provides web application implementations for PC and mobile, users can create high-quality images directly through the user interface.
Bing-Kun Bao, Yefei Sheng, Jie Wang 0061, Sisi You
AAAI5
2025 Spatial-Temporal Prior Knowledge Guidance for Long-term Action Anticipation
abstract
For long-term action anticipation (LAA), the primary focus is on understanding the observed video content and anticipating future actions, including both the names of upcoming actions and their corresponding durations. This requires the model to fully grasp the patterns of action transitions. To facilitate this, we employ spatial-temporal prior knowledge to guide the LAA model in capturing these transition patterns, which is referred to as explicit learning. Additionally, we use a transformer structure which incorporates a parallel decoding mechanism in which mitigates error accumulation and can be regarded as implicit learning. Consequently, we propose a novel model that integrates explicit and implicit learning approaches, combining the advantages of both. On the benchmarks for long-term action anticipation, our method achieves state-of-the-art results on the 50Salads and Breakfast.
Yiming Li 0008, Miao Ji, Sisi You, Bing-Kun Bao
ICME3
2025 SCVBench: A Benchmark with Multi-turn Dialogues for Story-Centric Video Understanding
abstract
Video understanding seeks to enable machines to interpret visual content across three levels: action, event, and story. Existing models are limited in their ability to perform high-level long-term story understanding, due to (1) the oversimplified treatment of temporal information and (2) the training bias introduced by action/event-centric datasets. To address this, we introduce SCVBench, a novel benchmark for story-centric video understanding. SCVBench evaluates LVLMs through an event ordering task decomposed into sub-questions leading to a final question, quantitatively measuring historical dialogue exploration. We collected 1,253 final questions and 6,027 sub-question pairs from 925 videos, constructing continuous multi-turn dialogues. Experimental results show that while closed-source GPT-4o outperforms other models, most open-source LVLMs struggle with story-centric video understanding. Additionally, our StoryCoT model significantly surpasses open-source LVLMs on SCVBench. SCVBench aims to advance research by comprehensively analyzing LVLMs' temporal reasoning and comprehension capabilities. Code can be accessed at https://github.com/yuanrr/SCVBench.
Sisi You, Bing-Kun Bao
IJCAI1
2025 DToMA: Training-free Dynamic Token MAnipulation for Long Video Understanding
abstract
Video Large Language Models (VideoLLMs) often require thousands of visual tokens to process long videos, leading to substantial computational costs, further exacerbated by visual token inefficiency. Existing token reduction and alternative video representation methods improve efficiency but often compromise comprehension abilities. In this work, we analyze the reasoning processes of VideoLLMs in multi-choice VideoQA task, identifying three reasoning stages—shallow, intermediate, and deep stages—that closely mimic human cognitive processing. Our analysis reveals specific inefficiencies at each stage: in shallow layers, VideoLLMs attempt to memorize all video details without prioritizing relevant content; in intermediate layers, models fail to re-examine uncertain content dynamically; and in deep layers, they continue processing video even when sufficiently confident. To bridge this gap, we propose DToMA, a training-free Dynamic Token MAnipulation method inspired by human adjustment mechanisms in three aspects: 1) Text-guided keyframe-aware reorganization to prioritize keyframes and reduce redundancy, 2) Uncertainty-based visual injection to revisit content dynamically, and 3) Early-exit pruning to halt visual tokens when confident. Experiments on 6 long video understanding benchmarks show that DToMA enhances both efficiency and comprehension, outperforming state-of-the-art methods and generalizing well across 3 VideoLLM architectures and sizes. Code is available at https://github.com/yuanrr/DToMA.
Sisi You, Bing-Kun Bao
IJCAI2
2024 Dynamic Scene Graph Generation with Unified Temporal Modeling
abstract
Dynamic scene graph generation requires understanding the spatial information intra-frame and temporal information between different frames. Existing methods utilize implicit and explicit modeling algorithms to capture temporal information and correlations by designing network architectures or incorporating prior knowledge. However, they exclusively rely on relationship evolution patterns within distinct post-processing modules that are independent of the temporal encoder, leading to solely modifying the results of relationship prediction and difficulty fully harnessing the temporal cues inherent in the video. To address the above challenge, we propose a Unified Temporal Modeling (UTM) that can integrate temporal encoding and temporal correlation modeling. We leverage the relationship evolution patterns to model temporal correlations that can be adapted to capture more relevant temporal cues during temporal encoding. Additionally, our model can be applied to existing image-based scene graph generation methods, extending their capabilities to video tasks. Extensive experiments on the Action Genome dataset demonstrate the robustness of UTM.
Sisi You, Bing-Kun Bao
ICME1
2024 Two-Stage Reasoning Network with Modality Decomposition for Text VQA
Shengrong Ling, Sisi You, Bing-Kun Bao
MMM (3)2
2024 Multi-object Tracking with Spatial-Temporal Tracklet Association
abstract
Recently, the tracking-by-detection methods have achieved excellent performance in Multi-Object Tracking (MOT), which focuses on obtaining a robust feature for each object and generating tracklets based on feature similarity. However, they are confronted with two issues: (1) unstable features in short-term occlusion and (2) insufficient matching in long-term occlusion. Specifically, the unstable feature is caused by the appearance variation under occlusion, and the association with the current unstable feature will lead to insufficient matching in long-term occlusion. To address the above issues, we propose a two-stage tracklet-level association method, Spatial-Temporal Tracklet Association (STTA), to effectively combine spatial-temporal context between feature extraction and data association. In the first stage, we propose the Tracklet-guided Spatial-Temporal Attention network (TSTA) to generate robust and stable features. Specifically, TSTA captures spatial-temporal context to obtain the most salient regions between the current and previous clips. In the second stage, we design the Bi-Tracklet Spatial-Temporal association (BTST) module to fully exploit the spatial-temporal context in data association. Specifically, we leverage BTST to merge different tracklets into long-term trajectories by jointly learning visual feature and spatial-temporal context and designing a bidirectional interpolation to recover the missed objects between matched tracklets. Extensive experiments of public and private detections on four benchmarks demonstrate the robustness of STTA. Furthermore, the proposed method is a model-agnostic method, which can be plugged and played with existing methods to boost their performance, e.g., obtain 11.0%, 10.1%, 2.9%, 3.2%, and 7.8% improvement on IDF1 in the MOT16 validation dataset for Tracktor, CenterTrack, Deepsort, JDE, and CTracker, respectively.
Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Incremental Audio-Visual Fusion for Person Recognition in Earthquake Scene
abstract
Earthquakes have a profound impact on social harmony and property, resulting in damage to buildings and infrastructure. Effective earthquake rescue efforts require rapid and accurate determination of whether any survivors are trapped in the rubble of collapsed buildings. While deep learning algorithms can enhance the speed of rescue operations using single-modal data (either visual or audio), they are confronted with two primary challenges: insufficient information provided by single-modal data and catastrophic forgetting. In particular, the complexity of earthquake scenes means that single-modal features may not provide adequate information. Additionally, catastrophic forgetting occurs when the model loses the information learned in a previous task after training on subsequent tasks, due to non-stationary data distributions in changing earthquake scenes. To address these challenges, we propose an innovative approach that utilizes an incremental audio-visual fusion model for person recognition in earthquake rescue scenarios. Firstly, we leverage a cross-modal hybrid attention network to capture discriminative temporal context embedding, which uses self-attention and cross-modal attention mechanisms to combine multi-modality information, enhancing the accuracy and reliability of person recognition. Secondly, an incremental learning model is proposed to overcome catastrophic forgetting, which includes elastic weight consolidation and feature replay modules. Specifically, the elastic weight consolidation module slows down learning on certain weights based on their importance to previously learned tasks. The feature replay module reviews the learned knowledge by reusing the features conserved from the previous task, thus preventing catastrophic forgetting in dynamic environments. To validate the proposed algorithm, we collected the Audio-Visual Earthquake Person Recognition (AVEPR) dataset from earthquake films and real scenes. Furthermore, the proposed method gets 85.41% accuracy while learning the 10th new task, which demonstrates the effectiveness of the proposed method and highlights its potential to significantly improve earthquake rescue efforts.
Sisi You, Yukun Zuo, Hantao Yao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Unbiased Feature Learning with Causal Intervention for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match individuals across different modalities. Existing methods can learn class-separable features but still struggle with modality gaps within class due to the modality-specific information, which is discriminative in one modality but not present in another (e.g., a black striped shirt). The presence of the interfering information creates a spurious correlation with the class label, which hinders alignment across modalities. To this end, we propose an Unbiased feature learning method based on Causal inTervention for VI-ReID from three aspects. Firstly, through the proposed structural causal graph, we demonstrate that modality-specific information acts as a confounder that restricts the intra-class feature alignment. Secondly, we propose a causal intervention method to remove the confounder using an effective approximation of backdoor adjustment, which involves adjusting the spurious correlation between features and labels. Thirdly, we incorporate the proposed approximation method into the basic VI-ReID model. Specifically, the confounder can be removed by adjusting the extracted features with a set of weighted pre-trained class prototypes from different modalities, where the weight is adapted based on the features. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that our method outperforms state-of-the-art methods. Code is available at https://github.com/NJUPT-MCC/UCT .
Sisi You, Bing-Kun Bao
ACM Trans. Multim. Comput. Commun. Appl.3
2023 UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement
abstract
Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM.
Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu
CVPR1
2023 Self-PT: Adaptive Self-Prompt Tuning for Low-Resource Visual Question Answering
abstract
Pretraining and finetuning large vision-language models (VLMs) have achieved remarkable success in visual question answering (VQA). However, finetuning VLMs requires heavy computation, expensive storage costs, and is prone to overfitting for VQA in low-resource settings. Existing prompt tuning methods have reduced the number of tunable parameters, but they cannot capture valid context-aware information during prompt encoding, resulting in 1) poor generalization of unseen answers and 2) lower improvements with more parameters. To address these issues, we propose a prompt tuning method for low-resource VQA named Adaptive Self-Prompt Tuning (Self-PT), which utilizes representations of question-image pairs as conditions to obtain context-aware prompts. To enhance the generalization of unseen answers, Self-PT uses dynamic instance-level prompts to avoid overfitting the correlations between static prompts and seen answers observed during training. To reduce parameters, we utilize hyper-networks and low-rank parameter factorization to make Self-PT more flexible and efficient. The hyper-network decouples the number of parameters and prompt length to generate flexible-length prompts by the fixed number of parameters. While the low-rank parameter factorization decomposes and reparameterizes the weights of the prompt encoder into a low-rank subspace for better parameter efficiency. Experiments conducted on VQA v2, GQA, and OK-VQA with different low-resource settings show that our Self-PT outperforms the state-of-the-art parameter-efficient methods, especially in lower-shot settings, e.g., 6% average improvements cross three datasets in 16-shot. Code is available at https://github.com/NJUPT-MCC/Self-PT.
Sisi You, Bing-Kun Bao
ACM Multimedia2
2023 Rescue decision via Earthquake Disaster Knowledge Graph reasoning
Yifan Jiao, Sisi You
Multim. Syst.2
2022 Multi-Object Tracking With Spatial-Temporal Topology-Based Detector
abstract
Multi-object tracking is a challenging task due to the occlusion of different targets. Existing methods focus on inferring a robust and discriminative feature for data association based on the targets generated by the existing detector. Unlike existing methods that consider each target independently during generating the trajectories, we propose a novel Spatial-Temporal Topology-based Detector (STTD) algorithm that treats the target and its nearest neighbors as a cluster and introduces a topology structure to describe the dynamics of moving targets belonging to the same cluster. With the public detections and the tracked objects in the previous frame, STTD firstly refines them by regression of detector to obtain the candidate proposals in the current frame. After that, the temporal topology constraint is proposed to recover the missed objects by considering the continuity and consistency of the topological structure. Based on the assumption that the targets belonging to the same topology should have a consistent characteristic, the spatial topology constraint is proposed to remove the inaccurate targets. Then we can obtain new candidate objects and construct the cost matrix used for data association. The evaluations on three MOTChallenge benchmarks verify the effectiveness of the proposed method.
Sisi You, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2021 Multi-Target Multi-Camera Tracking With Optical-Based Pose Association
abstract
Multi-target multi-camera tracking (MTMCT) targets to generate trajectories of the object that appeared under multiple cameras automatically. MTMCT can be treated as a combination of intra-camera tracking and cross-camera tracking. The existing work only employs the global description to perform the tracklet generating. However, the global description cannot model the local similarity between targets, leading to existing methods not to be robust to occlusion and fast motion. To handle the mentioned problem, we propose an online Optical-based Pose Association (OPA) for multi-target multi-camera tracking. The proposed method utilizes local pose matching to solve the occlusion problem, and applies optical flow to reduce the distance caused by fast motion. For optical-based pose association, we firstly employ OpenPose to generate human pose for each proposal. Then, we utilize the optical flow generated by PWC-Net to adjust the estimated pose for the previous frame. Finally, the modified Object Keypoint Similarity is used to compute the similarity between the pose of the current frame and adjusted pose in the prior frame. Once obtaining the optical-based pose similarity, we combine it with the visual and bounding box spatial similarities to generate the final similarity matrix, and apply the Kuhn-Munkras algorithm for data association. The experiments on the MTMCT and MOT datasets verify the rationality of using human pose information and prove the superiority of the proposed method.
Sisi You, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1