VLDB 2026 Research / reviewers in the wild / expert
Yaosi Hu
dblp:231/2696
· DBLP profile ↗
21ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-2784-6738ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoGenVD: A Motion-Centered Quality Assessment Benchmark for Text-to-Video Generation
Yingxue Zhang 0004, Zike Yang, Zihang Su, Yaosi Hu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | MotionPrior: Exploring Efficient Learning of Motion Concepts for Few-Shot Video GenerationabstractThe diffusion-based text-to-image generation has achieved remarkable progress and realistic content generation performance, greatly promoting the development in text-to-video generation. Although equipped with powerful image diffusion models, video generation modeling still requires massive labeled data and a high training resource cost. Recent, work has been focused on cost-effective video generation in a one-shot or few-shot manner based on the image diffusion model with minimum demand for video data and computing resources. However, these video generation models only support the generation of one single motion pattern/concept. This raises an important question: Can we improve generation freedom with a light training burden? In this paper, we explore a cost-effective video generation scheme for adaptive motion concepts by learning motion priors from a small set of video data. Specifically, we construct a learnable bank for motion concepts and propose the Dual-Semantic-guided Motion Attention module to locate the corresponding motion elements from the bank with the guidance of textual semantic and visual semantic. The extracted motion elements are inserted into video latents via lightweight motion injection layer, which is capable of integrating motion semantic effectively with much fewer parameters compared to the conventional temporal attention layer. In addition, we introduce a temporal-aware noise prior and an inter-frame consistency constraint to strengthen the learning of temporal dependency and improve video smoothness. Extensive experiments validate that the proposed method can learn motion priors adaptively from a small set of training videos to generate smooth videos that involve either single or multiple motion concepts. The results demonstrate that the proposed scheme achieves superior performance compared to existing few-shot video generation methods and even some large-scale video generation models. More information and results are available at https://youncy-hu.github.io/motionprior/. Yaosi Hu, Chang Wen Chen |
IEEE Trans. Image Process. | 1 |
| 2025 | SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject ControlabstractAutonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field. Binyuan Huang, Yuqing Wen, Yaosi Hu, Yingfei Liu, Fan Jia 0006, Weixin Mao, Tiancai Wang, Chi Zhang 0026, Chang Wen Chen, Zhenzhong Chen 0001, Xiangyu Zhang 0005 |
AAAI | 4 |
| 2025 | TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationabstractText-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and the textual description. (ii) how to improve the subjective quality of generated videos. To tackle the above challenges, we propose a new diffusion-based TI2V framework, termed TIV-Diffusion, via object-centric textual-visual alignment, intending to achieve precise control and high-quality video generation based on textual-described motion for different objects. Concretely, we enable our TIV-Diffuion model to perceive the textual-described objects and their motion trajectory by incorporating the fused textual and visual knowledge through scale-offset modulation. Moreover, to mitigate the problems of object disappearance and misaligned objects and motion, we introduce an object-centric textual-visual alignment module, which reduces the risk of misaligned objects/motion by decoupling the objects in the reference image and aligning textual features with each object individually. Based on the above innovations, our TIV-Diffusion achieves state-of-the-art high-quality video generation compared with existing TI2V methods. Xingrui Wang, Xin Li 0082, Yaosi Hu, Hanxin Zhu, Chen Hou, Cuiling Lan, Zhibo Chen 0001 |
AAAI | 3 |
| 2025 | LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
Yaosi Hu, Zhenzhong Chen 0001, Chong Luo 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Remote Sensing Semantic Segmentation Quality Assessment Based on Vision Language ModelabstractVariations in scene complexity and image quality across remote sensing images lead to inconsistent performance when applying pre-trained semantic segmentation models. To ensure quality control for downstream tasks, quality assessment on interpretation results becomes a challenging task, primarily due to the limited availability of reference annotations. To address this issue, we propose RS-SQA, a remote sensing semantic segmentation quality assessment model based on the vision language model (VLM) for images without ground truth. RS-SQA uses a dual-branch architecture that leverages a RS VLM for semantic understanding and utilizes intermediate features from segmentation methods to extract implicit information. We fuse these two feature streams via a semantic-guided module, yielding accurate quality predictions. Specifically, we introduce CLIP-RS, a large-scale pre-trained VLM trained on 10M remote sensing images with purified captions. Feature visualizations confirm that CLIP-RS can effectively differentiate between various levels of segmentation quality. To facilitate the development of RS semantic segmentation quality assessment, we present RS-SQED, a new dedicated dataset sampled from 4 major RS semantic segmentation datasets and annotated with segmentation accuracy scores from 8 representative segmentation methods. Experiments on the established dataset demonstrate that RS-SQA outperforms state-of-the-art quality assessment models and aids in selecting the most suitable pre-trained models for remote sensing semantic segmentation by comparing predicted accuracies across different models. Huiying Shi, Zhihong Tan, Hongchen Wei, Yaosi Hu, Yingxue Zhang 0004, Zhenzhong Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Memory-guided representation matching for unsupervised video anomaly detectionabstractRecent works on Video Anomaly Detection (VAD) have made advancements in the unsupervised setting, known as Unsupervised VAD (UVAD), which brings it closer to practical applications. Unlike the classic VAD task that requires a clean training set with only normal events, UVAD aims to identify abnormal frames without any labeled normal/abnormal training data . Many existing UVAD methods employ handcrafted surrogate tasks, such as frame reconstruction, to address this challenge. However, we argue that these surrogate tasks are sub-optimal solutions, inconsistent with the essence of anomaly detection. In this paper, we propose a novel approach for UVAD that directly detects anomalies based on similarities between events in videos. Our method generates representations for events while simultaneously capturing prototypical normality patterns, and detects anomalies based on whether an event’s representation matches the captured patterns. The proposed model comprises a memory module to capture normality patterns, and a representation learning network to obtain representations matching the memory module for normal events. A pseudo-label generation module as well as an anomalous event generation module for negative learning are further designed to assist the model to work under the strictly unsupervised setting. Experimental results demonstrate that the proposed method outperforms existing UVAD methods and achieves competitive performance compared with classic VAD methods. Yiran Tao, Yaosi Hu, Zhenzhong Chen 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | A Benchmark for Controllable Text -Image-to-Video GenerationabstractAutomatic video generation is a challenging research topic, attracting interests from different perspectives, including Image-to-Video generation (I2V), Video-to-Video generation (V2V), and Text-to-Video generation (T2V). To pursue more controllable and fine-grained video generation, a novel video generation task, named Text-Image-to-Video generation (TI2V), and a corresponding baseline solution, named Motion Anchor-based video Generator (MAGE), were proposed. However, two other factors, namely clean datasets and reliable evaluation metrics, also play important roles in the success of the TI2V task. In this article, we present a complete benchmark for the TI2V task which includes synthetic video-text paired datasets, a baseline method, and two evaluation metrics. More specifically: (1) Two versions of synthetic datasets are built based on CATER containing rich combinations of objects and actions, as well as the resulting changes of brightness and shadow. We also provide both explicit and ambiguous text descriptions to support deterministic and diverse video generation, respectively. (2) A refined version of MAGE, dubbed MAGE+, is proposed with an innovative motion anchor structure to store appearance-motion aligned representation, which can be further injected with explicit condition and implicit randomness to model the uncertainty in data distribution. (3) To evaluate the quality of generated video especially given ambiguous description, we introduce action precision and referring expression precision to assess the quality of motion based on captioning-and-matching method. Experiments conducted on proposed datasets, as well as relevant datasets, verify the effectiveness of our baseline and show appealing potentials of TI2V task. Yaosi Hu, Chong Luo 0001, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | A Lightweight No-reference Video Quality Assessment MethodabstractRecently, quality assessment for user-generated content (UGC) videos has become a challenging task due to the absence of reference videos and the presence of complex distortions. Prior methods has highlighted the effectiveness of semantic features for quality assessment. However, these models are incapable for real-time prediction and efficient computation in practical applications. In this paper, we design a lightweight no-reference video quality assessment model leveraging pretrained lightweight network for semantic understanding and utilizing a low-level CNN for distortion features. The temporal features and spatial features are extracted respectively in the semantic and low-level scales, and then they are multiplied and integrated to obtain the video-level score. Experiments on UGC video quality databases show that our proposed model achieves comparable accuracy to state-of-the-art benchmarks while providing real-time performance on GPU. Huiying Shi, Yaosi Hu, Yingxue Zhang 0004, Zhenzhong Chen 0001 |
VCIP | 2 |
| 2023 | Multiple visual relationship forecasting and arrangement in videos
Wanping Ouyang, Yaosi Hu, Yangjun Ou, Zhenzhong Chen 0001 |
Neurocomputing | 2 |
| 2022 | Make It Move: Controllable Image-to-Video Generation with Text DescriptionsabstractGenerating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video generation (TI2V), is proposed. With both controllable appearance and motion, TI2V aims at generating videos from a static image and a text description. The key challenges of TI2V task lie both in aligning appearance and motion from different modalities, and in handling uncertainty in text descriptions. To address these challenges, we propose a Motion Anchor-based video GEnerator (MAGE) with an innovative motion anchor (MA) structure to store appearance-motion aligned representation. To model the uncertainty and increase the diversity, it further allows the injection of explicit condition and implicit randomness. Through three-dimensional axial transformers, MA is interacted with given image to generate next frames recursively with satisfying controllability and diversity. Accompanying the new task, we build two new video-text paired datasets based on MNIST and CATER for evaluation. Experiments conducted on these datasets verify the effectiveness of MAGE and show appealing potentials of TI2V task. Datasets are available at https://github.com/Youncy-Hu/MAGE. Yaosi Hu, Chong Luo 0001, Zhenzhong Chen 0001 |
CVPR | 1 |
| 2022 | Video Quality Assessment based on Quality Aggregation NetworksabstractA reliable video quality assessment (VQA) algorithm is essential for evaluating and optimizing video processing pipelines. In this paper, we propose a quality aggregation network (QAN) for full-reference VQA, which models the characteristics of human visual perception of video quality in both spatial and temporal domain. The proposed QAN is composed of two mod-ules, the spatial quality aggregation (SQA) network and the tem-poral quality aggregation (TQA) network. Specifically, the SQA network models the quality of video frames using 3D CNN, taking both spatial and temporal masking effects into consideration for the modeling of the perception of human visual system (HVS). In the TQA network, considering the memory effect of HVS facing the temporal variation of frame-level quality, an LSTM-based temporal quality pooling network is proposed to capture the nonlinearities and temporal dependencies involved in the process of quality evaluation. According to the experimental results on two well-established VQA databases, the proposed model could outperform the state-of-the-art metrics. The code of the proposed method is available at: https://github.com/lorenzowu/QAN. Wei Wo, Yingxue Zhang 0004, Yaosi Hu, Zhenzhong Chen 0001, Shan Liu 0001 |
VCIP | 3 |
| 2022 | Decomposing style, content, and motion for videos
Yaosi Hu, Dacheng Yin, Yuwang Wang, Zhenzhong Chen 0001, Chong Luo 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Predicate Correlation Learning for Scene Graph GenerationabstractFor a typical Scene Graph Generation (SGG) method in image understanding, there usually exists a large gap in the performance of the predicates’ head classes and tail classes. This phenomenon is mainly caused by the semantic overlap between different predicates as well as the long-tailed data distribution. In this paper, a Predicate Correlation Learning (PCL) method for SGG is proposed to address the above problems by taking the correlation between predicates into consideration. To measure the semantic overlap between highly correlated predicate classes, a Predicate Correlation Matrix (PCM) is defined to quantify the relationship between predicate pairs, which is dynamically updated to remove the matrix’s long-tailed bias. In addition, PCM is integrated into a predicate correlation loss function (LPC) to reduce discouraging gradients of unannotated classes. The proposed method is evaluated on several benchmarks, where the performance of the tail classes is significantly improved when built on existing methods. Leitian Tao, Li Mi, Nannan Li 0004, Xianhang Cheng, Yaosi Hu, Zhenzhong Chen 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 15 |
| 2021 | Learn to Look Around: Deep Reinforcement Learning Agent for Video Saliency PredictionabstractIn the video saliency prediction task, one of the key issues is the utilization of temporal contextual information of keyframes. In this paper, a deep reinforcement learning agent for video saliency prediction is proposed, designed to look around adjacent frames and adaptively generate a salient contextual window that contains the most correlated information of keyframe for saliency prediction. More specifically, an action set step by step decides whether to expand the window, meanwhile a state set and reward function evaluate the effectiveness of the current window. The deep Q-learning algorithm is followed to train the agent to learn a policy to achieve its goal. The proposed agent can be regarded as plug-and-play which is compatible with generic video saliency prediction models. Experimental results on various datasets demonstrate that our method can achieve an advanced performance. Yiran Tao, Yaosi Hu, Zhenzhong Chen 0001 |
VCIP | 2 |
| 2021 | MAPS: Joint Multimodal Attention and POS Sequence Generation for Video CaptioningabstractVideo captioning is considered to be challenging due to the combination of video understanding and text generation. Recent progress in video captioning has been made mainly using methods of visual feature extraction and sequential learning. However, the syntax structure and semantic consistency of generated captions are not fully explored. Thus, in our work, we propose a novel multimodal attention based framework with Part-of-Speech (POS) sequence guidance to generate more accu-rate video captions. In general, the word sequence generation and POS sequence prediction are hierarchically jointly modeled in the framework. Specifically, different modalities including visual, motion, object and syntactic features are adaptively weighted and fused with the POS guided attention mechanism when computing the probability distributions of prediction words. Experimental results on two benchmark datasets, i.e. MSVD and MSR-VTT, demonstrate that the proposed method can not only fully exploit the information from video and text content, but also focus on the decisive feature modality when generating a word with a certain POS type. Thus, our approach boosts the video captioning performance as well as generating idiomatic captions. Cong Zou, Yaosi Hu, Zhenzhong Chen 0001, Shan Liu 0001 |
VCIP | 3 |
| 2020 | A Multimodal Variational Encoder-Decoder Framework for Micro-video Popularity PredictionabstractPredicting the popularity of a micro-video is a challenging task, due to a number of factors impacting the distribution such as the diversity of the video content and user interests, complex online interactions, etc. In this paper, we propose a multimodal variational encoder-decoder (MMVED) framework that considers the uncertain factors as the randomness for the mapping from the multimodal features to the popularity. Specifically, the MMVED first encodes features from multiple modalities in the observation space into latent representations and learns their probability distributions based on variational inference, where only relevant features in the input modalities can be extracted into the latent representations. Then, the modality-specific hidden representations are fused through Bayesian reasoning such that the complementary information from all modalities is well utilized. Finally, a temporal decoder implemented as a recurrent neural network is designed to predict the popularity sequence of a certain micro-video. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed model in the micro-video popularity prediction task. Jiayi Xie, Yaochen Zhu, Jing Yi, Yaosi Hu, Hongyi Liu 0003, Zhenzhong Chen 0001 |
WWW | 6 |
| 2020 | Exploiting the local temporal information for video captioning
Li Mi, Yaosi Hu, Zhenzhong Chen 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Two-Stream Refinement Network for RGB-D Saliency DetectionabstractIn this paper, we propose a two-stream refinement network for RGB-D saliency detection. A fusion refinement module is designed to fuse output features from different resolution and modals. The structure information from depth helps distinguish between foreground and background and the lower level features with higher resolution can be adopted to refine the boundary of detected targets. The proposed model predicts high-resolution saliency map and then use a propagation-based module to further refine object boundary. Experimental results demonstrate that the proposed method performs well against to the state of the art methods on the recent RGB-D salient object detection dataset. Di Liu 0007, Yaosi Hu, Kao Zhang, Zhenzhong Chen 0001 |
ICIP | 2 |
| 2019 | Hierarchical Global-Local Temporal Modeling for Video CaptioningabstractIn this paper, a Hierarchical Temporal Model (HTM) is proposed for the video captioning task, based on exploring the global and local temporal structure to better recognize fine-grained objects and actions. In our HTM, the encoder and decoder are hierarchically aligned according to different levels of features. The encoder applies two LSTM layers to construct temporal structures at both frame-level and object-level where the attention mechanism is applied to locate objects of interest, and the decoder uses corresponding LSTM layers to extract pivotal features from global to local through multi-level attention mechanism. Moreover, the local temporal structure is constructed implicitly from candidate object-oriented features under the guidance of global temporal-spatial representation, that could generate more accurate descriptions in handling shot-switching problems. Experiments on the widely used Microsoft Video Description Corpus (MSVD) and Charades datasets demonstrate the effectiveness of our proposed approach when compared to the state-of-the-art methods. Yaosi Hu, Zhenzhong Chen 0001, Zhengjun Zha, Feng Wu 0001 |
ACM Multimedia | 1 |