Ruixiang Zhao

dblp:257/9470 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Co-teaching for Unsupervised Domain Expansion
Hailan Lin, Qijie Wei, Kaibin Tian, Ruixiang Zhao, Xirong Li 0001
MMM (1)4
2026 ASR-Enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
abstract
E-commerce is increasinglymultimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We proposeASR-enhancedMultimodalProduct Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text,AMPereuses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness ofAMPerein obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval.
Ruixiang Zhao, Jian Jia, Yan Li 0043, Xuehan Bai, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xirong Li 0001
IEEE Trans. Multim.1
2025 Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video Retrieval
abstract
The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers from low efficiency. In this study, we explore a new Hybrid-Tower framework that can hybridize the advantages of the Two-Tower and Single-Tower framework, achieving high effectiveness and efficiency simultaneously. We propose a novel hybrid method, Fine-grained Pseudo-query Interaction and Generation for T2VR, ie, PIG, which includes a new pseudo-query generator designed to generate a pseudo-query for each video. This enables the video feature and the textual features of pseudo-query to interact in a fine-grained manner, similar to the Single-Tower approaches to hold high effectiveness, even before the real textual query is received. Simultaneously, our method introduces no additional storage or computational overhead compared to the Two-Tower framework during the inference stage, thus maintaining high efficiency. Extensive experiments on five commonly used text-video retrieval benchmarks demonstrate that our method achieves a significant improvement over the baseline, with an increase of $1.6\% \sim 3.9\%$ in R@1. Furthermore, our method matches the efficiency of Two-Tower models while achieving near state-of-the-art performance, highlighting the advantages of the Hybrid-Tower framework.
Bangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun, Zhanhui Kang, Gang Yang 0001, Xirong Li 0001
ICCV3
2025 Multi-Object Sketch Animation by Scene Decomposition and Motion Planning
abstract
Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we identify two major challenges of transitioning from single-object to multi-object sketch animation: object-aware motion modeling and complex motion optimization. For multi-object sketch animation, we propose MoSketch based on iterative optimization through Score Distillation Sampling (SDS) and thus animating a multi-object sketch in a training-data free manner. To tackle the two challenges in a divide-and-conquer strategy, MoSketch has four novel modules, i.e., LLM-based scene decomposition, LLM-based motion planning, multi-grained motion refinement, and compositional SDS. Extensive qualitative and quantitative experiments demonstrate the superiority of our method over existing sketch animation approaches. MoSketch takes a pioneering step towards multi-object sketch animation, opening new avenues for future research and applications.
Zijie Xin, Yuhan Fu, Ruixiang Zhao, Bangxiang Lan, Xirong Li 0001
ICCV4
2025 Op-Occ-Net: Neural Ordinary Differential Equation Based 4D Occupancy Forecasting for Autonomous Vehicles
Junfeng Ding, Erxin Guo, Pei An, Jie Ma 0003, Ruixiang Zhao
PRCV (11)5
2024 Holistic Features are Almost Sufficient for Text-to-Video Retrieval
abstract
For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact, state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature interaction and matching, putting their scalability for large-scale T2VR applications into doubt. We propose TeachCLIP, enabling a CLIP4Clip based student network to learn from more advanced yet computationally intensive models. In order to create a learning channel to convey fine-grained cross-modal knowledge from a heavy model to the student, we add to CLIP4Clip a simple Attentionalframe-Feature Aggregation (AFA) block, which by design adds no extra storage /computation overhead at the retrieval stage. Frame-text relevance scores calculated by the teacher network are used as soft labels to supervise the attentive weights produced by AFA. Extensive experiments on multiple public datasets justify the viability of the proposed method. TeachCLIP has the same efficiency and compact-ness as CLIP4Clip, yet has near-SOTA effectiveness.
Kaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan, Xirong Li 0001
CVPR2
2024 How SAM helps Unsupervised Video Object Segmentation?
abstract
As an emerging vision foundation model, Segment Anything Model (SAM) has been successfully applied to Video Object Segmentation (VOS). However, previous methods rely on first frame’s mask or manual interaction, which belong to semi-supervised or interactive video object segmentation. The potential of SAM in Unsupervised Video Object Segmentation (UVOS) remains to be explored. In this work, we propose a two-stage training framework to explore how SAM helps UVOS. In Stage 1, utilizing SAM powerful feature extraction ability, we only train a lightweight decoder with feature aggregation module. This stage employs relaxed flow reconstruction as an unsupervised proxy task for object discovery. In Stage 2, based on SAM promptability, we design a refinement process that involves rough correction based on in-clip temporal consistency and heuristic refinement guided by Prompt-closure, to generate refined results from the coarse mask produced by Stage 1. Finally, update the segmentation head using refined results along with corresponding confidence. The proposed method demonstrates competitive performance on three common UVOS datasets (DAVIS2016, SegTrackv2, FBMS59) with higher accuracy, faster convergence, lower training cost, validating the effectiveness of our approach.
Jiahe Yue, Runchu Zhang, Zhe Zhang 0037, Ruixiang Zhao, Wu Lv, Jie Ma 0003
IJCNN4
2023 ChinaOpen: A Dataset for Open-world Multimodal Learning
abstract
This paper introduces ChinaOpen, a dataset sourced from Bilibili, a popular Chinese video-sharing website, for open-world multimodal learning. While the state-of-the-art multimodal learning networks have shown impressive performance in automated video annotation and cross-modal video retrieval, their training and evaluation are primarily conducted on YouTube videos with English text. Their effectiveness on Chinese data remains to be verified. In order to support multimodal learning in the new context, we construct ChinaOpen-50k, a webly annotated training set of 50k Bilibili videos associated with user-generated titles and tags. Both text-based and content-based data cleaning are performed to remove low-quality videos in advance. For a multi-faceted evaluation, we build ChinaOpen-1k, a manually labeled test set of 1k videos. Each test video is accompanied with a manually checked user title and a manually written caption. Besides, each video is manually tagged to describe objects / actions / scenes shown in the visual content. The original user tags are also manually checked. Moreover, with all the Chinese text translated into English, ChinaOpen-1k is also suited for evaluating models trained on English data. In addition to ChinaOpen, we propose Generative Video-to-text Transformer (GVT) for Chinese video captioning. We conduct an extensive evaluation of the state-of-the-art single-task / multi-task models on the new dataset, resulting in a number of novel findings and insights.
Aozhu Chen, Chengbo Dong, Kaibin Tian, Ruixiang Zhao, Xun Liang 0001, Zhanhui Kang, Xirong Li 0001
ACM Multimedia5