Jiaxi Gu

dblp:207/0185 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 4 · 4 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 LSTD: Long Short-Term Temporal Diffusion for Video Generation
abstract
Recently, text-driven video generation has achieved tremendous progress. However, existing methods neglect the contexts of long short-term frames in the video, thereby compromising temporal consistency. They also encounter challenges of heavy memory costs due to the use of the standard temporal attention mechanism and misalignment between training videos and captions. Additionally, previous approaches for long video generation are flawed because they are hard to ensure content diversity and consistency. To alleviate these issues, we propose a novel Long Short-term Temporal Diffusion (LSTD) model to generate videos with superior temporal consistency. We introduce two novel temporal modules,i.e., the Short-term Temporal Convolution and the Long-term Temporal Attention. The former can learn short-term features with a shallow structure, and the latter concentrates on long-term information of complex motion with a new memory-efficient attention mechanism. The combination of the two modules can ensure the temporal consistency of the generated videos. Furthermore, a novel inference method for long video generation is also proposed, which can iteratively generate hundreds of video frames. Experimental results on UCF-101, MSR-VTT, and two long video benchmarks prove that our method achieves superior zero-shot inference performance even when the size of the training data is reduced by 26.5 times.
Jiaxi Gu, Shicong Wang, Xing Zhang 0013, Zuxuan Wu, Hang Xu 0004, Yu-Gang Jiang 0001
IEEE Trans. Multim.2
2025 EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation
abstract
The diffusion model is widely leveraged for either controllable video generation or video interpolation. As each field has its task-specific problems, it is difficult to merely develop a single model for completing both tasks simultaneously. Moreover, most existing works only support image conditions and necessitate redesigning the model structure to accommodate other types of conditions. Even so, they still face frame flickering issues when using the image as the condition due to the strong alignment of image pixels. To tackle these problems, in this work, we are the first to propose a unified diffusion framework, EasyControl, for both tasks of controllable video generation and interpolation with different types of conditions. The proposed EasyControl introduces a condition adapter to extract the condition features, which is then injected into an interchangeable fundamental text-to-video model to guide the video generation. To alleviate frame flicker problems, we propose a module named VideoInit to integrate the low-frequency band of input condition images, ensuring smoother generation. Experimental results on four benchmarks suggest that our method outperforms the previous methods on each task.
Jiaxi Gu, Panwen Hu, Yuanfan Guo, Xiaodan Liang
ICASSP2
2025 DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
abstract
Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. To tackle these problems, we propose a high-fidelity image-to-video generation method by devising a frame retention branch based on a pre-trained video diffusion model, named DreamVideo. Our DreamVideo perceives the reference image via convolution layers and concatenates the features with the noisy latents as model input. By this means, the details of the reference image can be preserved to the greatest extent. In addition, by incorporating the designed double-condition classifier-free guidance, DreamVideo can generate high-quality videos of different actions by providing varying prompt texts. We conduct comprehensive experiments on the public datasets, and both quantitative and qualitative results indicate that our method outperforms the state-of-the-art method.
Cong Wang 0018, Jiaxi Gu, Panwen Hu, Yuanfan Guo, Hang Xu 0004, Xiaodan Liang
ICASSP2
2025 MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
abstract
In recent years, while generative AI has advanced significantly in image generation, video generation continues to face challenges in controllability, length, and detail quality, which hinder its application. We present MimicMotion, a framework for generating high-quality human videos of arbitrary length using motion guidance. Our approach has several highlights. Firstly, we introduce confidence-aware pose guidance that ensures high frame quality and temporal smoothness. Secondly, we introduce regional loss amplification based on pose confidence, which reduces image distortion in key regions. Lastly, we propose a progressive latent fusion strategy to generate long and smooth videos. Experiments demonstrate the effectiveness of our approach in producing high-quality human motion videos. Videos and comparisons are available at https://tencent.github.io/MimicMotion.
Jiaxi Gu, Li-Wen Wang, Junqi Cheng, Yuefeng Zhu, Fangyuan Zou
ICML2
2025 UniAdapter: All-in-One Control for Flexible Video Generation
abstract
Condition-based video generation aims to create video content based on given information that describes specific subjects. However, most existing works can only utilize a single condition to guide the denoising process, thereby limiting their applicability to specific scenarios. Although some works attempt to accommodate multiple conditions within one framework, they often require multiple encoders, leading to inefficiencies in integrating multi-condition features. In this work, we present a framework that, with the support of the proposed Unified Adapter (UniAdapter), enables simultaneous multi-condition control of video generation within a single model. To effectively merge these conditions, we propose a novel Probabilistic Multi-condition Concatenator (PMC) module, which employs a unified encoder to accommodate multiple conditions and concatenate condition features at the pixel level to achieve fine-grained control. Following the PMC module, we employ 2D down-sampling blocks to refine features for injection into the Video Diffusion Model (VDM). Moreover, our UniAdapter is designed to be model-agnostic and compatible with any U-Net-based VDM, offering a versatile solution for improving video generation quality. Experimental results on public benchmarks UCF-101 and MSR-VTT show that our method achieves superior results in both quantitative and qualitative evaluations.
Cong Wang 0018, Panwen Hu, Yuanfan Guo, Jiaxi Gu, Jianhua Han, Hang Xu 0004, Xiaodan Liang
IEEE Trans. Circuits Syst. Video Technol.5
2024 BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
abstract
Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various down-stream image synthesis tasks, such as controllable image generation and image editing, while downstream video synthesis tasks are less explored for several reasons. First, it requires huge memory and computation overhead to train a video generation foundation model. Even with video foundation models, additional costly training is still required for downstream video synthesis tasks. Second, although some works extend image diffusion models into videos in a training-free manner, temporal consistency cannot be well preserved. Finally, these adaption methods are specifically designed for one task and fail to generalize to different tasks. To mitigate these issues, we propose a training-free general-purpose video synthesis framework, coined as BIVDiff, via bridging specific image diffusion models and general text-to-video foundation diffusion models. Specifically, we first use a specific image diffusion model (e.g., ControlNet and Instruct Pix2Pix) for frame-wise video generation, then perform Mixed Inversion on the generated video, and finally input the inverted latents into the video diffusion models (e.g., VidRD and ZeroScope) for temporal smoothing. This decoupled framework enables flexible image model selection for different purposes with strong task generalization and high efficiency. To validate the effectiveness and general use of BIVDiff, we perform a wide range of video synthesis tasks, including controllable video generation, video editing, video inpainting, and outpainting.
Fengyuan Shi 0001, Jiaxi Gu, Songcen Xu, Limin Wang 0002
CVPR2
2024 MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and Editing
Jiaxi Gu, Xing Zhang 0013, Qingping Zheng, Zuxuan Wu, Hang Xu 0004, Yu-Gang Jiang 0001
ECCV (18)3
2024 Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
abstract
Latent Diffusion Models (LDMs) are renowned for their powerful capabilities in image and video synthesis. Yet, compared to text-to-image (T2I) editing, text-to-video (T2V) editing suffers from a lack of decent temporal consistency and structure, due to insufficient pre-training data, limited model editability, or extensive tuning costs. To address this gap, we propose FLDM (Fused Latent Diffusion Model), a training-free framework that achieves high-quality T2V editing by integrating various T2I and T2V LDMs. Specifically, FLDM utilizes a hyper-parameter with an update schedule to effectively fuse image and video latents during the denoising process. This paper is the first to reveal that T2I and T2V LDMs can complement each other in terms of structure and temporal consistency, ultimately generating high-quality videos. It is worth noting that FLDM can serve as a versatile plugin, applicable to off-the-shelf image and video LDMs, to significantly enhance the quality of video editing. Extensive quantitative and qualitative experiments on popular T2I and T2V LDMs demonstrate FLDM's superior editing quality than state-of-the-art T2V editing methods.
Xing Zhang 0013, Jiaxi Gu, Renjing Pei, Songcen Xu, Xingjun Ma, Hang Xu 0004, Zuxuan Wu
ACM Multimedia3
2023 PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval
abstract
Text-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video retrieval. However, due to the modality difference between videos and images, how to effectively adapt CLIP to the video domain is still underexplored. In this paper, we investigate this problem from two aspects. First, we enhance the transferred image encoder of CLIP for fine-grained video understanding in a seamless fashion. Second, we conduct fine-grained contrast between videos and texts from both model improvement and loss design. Particularly, we propose a fine-grained contrastive model equipped with parallel isomeric attention and dynamic routing, namely PIDRo, for text-video retrieval. The parallel isomeric attention module is used as the video encoder, which consists of two parallel branches modeling the spatial-temporal information of videos from both patch and frame levels. The dynamic routing module is constructed to enhance the text encoder of CLIP, generating informative word representations by distributing the fine-grained information to the related word tokens within a sentence. Such model design provides us with informative patch, frame and word representations. We then conduct token-wise interaction upon them. With the enhanced encoders and the token-wise loss, we are able to achieve finer-grained text-video alignment and more accurate retrieval. PIDRo obtains state-of-the-art performance over various text-video retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, DiDeMo and ActivityNet.
Peiyan Guan, Renjing Pei, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu 0004, Songcen Xu, Youliang Yan, Edmund Y. Lam
ICCV6
2022 Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
abstract
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_\text{ViT-L}$ achieves an average accuracy of 73.03%. For the image-text retrieval task, it achieves a mean recall of 71.6% on AIC-ICC which is 12.9% higher than WenLan 2.0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e.g., Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to https://wukong-dataset.github.io/wukong-dataset/.
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 0002, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang 0196, Xin Jiang 0002, Chunjing Xu, Hang Xu 0004
NeurIPS1
2020 Alohomora: Motion-Based Hotword Detection in Head-Mounted Displays
abstract
With the development of multimedia and computer graphics technologies, virtual reality (VR) is attracting more and more attention from both the academic communities and industrial companies. A head-mounted display (HMD) is the core equipment of VR. It envelops the entire sight of the wearer and reacts to some specific actions, mainly the head movement. Different from common video watching or game playing, VR poses the strict requirement of immersion so interaction methods need to be carefully designed. The hotword-based interaction as a typical hands-free method is very suitable for VR scenarios. However, the traditional hotword detection methods use a microphone to permit audio signal analysis. They not only incur significant recording overheads but are also susceptible to the surrounding noises. Instead of using the audio signals, we propose a motion-based hotword detection method called Alohomora. A multivariate time series (MTS) classification is formulated for processing the sensor data from multiple dimensions and types of motion sensors. We use a word extraction method for extracting and selecting patterns from MTS of motion data. Then, a classification model is trained using those discriminative patterns and finally the hotword can be detected in time. Alohomora is purely based on the motion sensors in HMDs without using any extra components such as microphone. As head tracking is always necessary in VR applications themselves, the overhead of Alohomora is nearly negligible. Finally, through extensive experiments, the final detection accuracy of Alohomora can exceed 90%.
Jiaxi Gu, Zhiwen Yu 0001, Kele Shen
IEEE Internet Things J.1
2020 Human-Machine Cooperative Video Anomaly Detection
abstract
It is still a challenge to detect anomalous events in video sequences in the field of computer vision due to heavy object occlusions, varying crowded densities and complex situations. To address this, we propose a novel human-machine cooperative approach which uses human feedback on anomaly confirmation to inform and enhance video anomaly detection. Specifically, we analyze the spatio-temporal characteristics of sequential frames of a video from the appearance and motion perspective from which spatial and temporal features are identified and extracted. We then develop a convolutional autoencoder neural network to compute an abnormal score based on reconstruction errors. In this process, a group of experts will provide human feedback to a certain proportion of classified frames to be incorporated into the model, and also the final judgment for the event anomalies for training and classification. The proposed approach is evaluated on 3 publicly available surveillance datasets, showing improved accuracy and competitive performance (93.7% AUC) with respect to the best performance (90.6% AUC) of the state-of-the-art approaches. The approach has not been previously seen to the best of our knowledge.
Fan Yang 0040, Zhiwen Yu 0001, Liming Chen 0001, Jiaxi Gu, Qingyang Li 0002, Bin Guo 0001
Proc. ACM Hum. Comput. Interact.4
2019 Traffic-Based Side-Channel Attack in Video Streaming
abstract
Video streaming takes up an increasing proportion of network traffic nowadays. Dynamic adaptive streaming over HTTP (DASH) becomes the de facto standard of video streaming and it is adopted by Youtube, Netflix, and so on. Despite of the popularity, network traffic during video streaming shows an identifiable pattern which brings threat to user privacy. In this paper, we propose a video identification method using network traffic while streaming. Though there is bitrate adaptation in DASH streaming, we observe that the video bitrate trend remains relatively stable because of the widely used variable bit-rate (VBR) encoding. Accordingly, we design a robust video feature extraction method for eavesdropped video streaming traffic. Meanwhile, we design a VBR-based video fingerprinting method for candidate video set which can be built using downloaded video files. Finally, we propose an efficient partial matching method for computing similarities between video fingerprints and streaming traces to derive video identities. We evaluate our attack method in different scenarios for various video content, segment lengths, and quality levels. The experimental results show that the identification accuracy can reach up to 90% using only three-minute continuous network traffic eavesdropping.
Jiaxi Gu, Jiliang Wang, Zhiwen Yu 0001, Kele Shen
IEEE/ACM Trans. Netw.1
2018 Walls Have Ears: Traffic-based Side-channel Attack in Video Streaming
abstract
Video streaming takes up an increasing proportion of network traffic nowadays. Dynamic Adaptive Streaming over HTTP (DASH) becomes the de facto standard of video streaming and it is adopted by Youtube, Netflix, etc. Despite of the popularity, network traffic during video streaming shows identifiable pattern which brings threat to user privacy. In this paper, we propose a video identification method using network traffic while streaming. Though there is bitrate adaptation in DASH streaming, we observe that the video bitrate trend remains relatively stable because of the widely used Variable Bit-Rate (VBR) encoding. Accordingly, we design a robust video feature extraction method for eavesdropped video streaming traffic. Meanwhile, we design a VBR based video fingerprinting method for candidate video set which can be built using downloaded video files. Finally, we propose an efficient partial matching method for computing similarities between video fingerprints and streaming traces to derive video identities. We evaluate our attack method in different scenarios for various video content, segment lengths and quality levels. The experimental results show that the identification accuracy can reach up to 90 % using only three-minute continuous network traffic eavesdropping.
Jiaxi Gu, Jiliang Wang, Zhiwen Yu 0001, Kele Shen
INFOCOM1
2018 NASR: NonAuditory Speech Recognition with Motion Sensors in Head-Mounted Displays
Jiaxi Gu, Kele Shen, Jiliang Wang, Zhiwen Yu 0001
WASA1
2017 PIC: Enable Large-Scale Privacy Preserving Content-Based Image Search on Cloud
abstract
Many cloud platforms emerge to meet urgent requirements for large-volume personal image store, sharing and search. Though most would agree that images contain rich sensitive information (e.g., people, location and event) and people's privacy concerns hinder their participation into untrusted services, today's cloud platforms provide little support for image privacy protection. Facing large-scale images from multiple users, it is extremely challenging for the cloud to maintain the index structure and schedule parallel computation without learning anything about the image content and indices. In this work, we introduce a novel system PIC: A Privacy-preserving Image search system on Cloud, which is a step towards feasible cloud services which provide secure content-based large-scale image search with fine-grained access control. Users can search on others' images if they are authorized by the image owners. Majority of the computationally intensive jobs are handled by the cloud, and a querier can now simply send the query and receive the result. Specially, to deal with massive images, we design our system suitable for distributed and parallel computation and introduce several optimizations to further expedite the search process. Our security analysis and prototype system evaluation results show that PIC successfully protects the image privacy at a low cost of computation and communication.
Lan Zhang 0002, Taeho Jung, Kebin Liu 0001, Xiang-Yang Li 0001, Jiaxi Gu, Yunhao Liu 0001
IEEE Trans. Parallel Distributed Syst.6