Yuanzhi Liang

dblp:193/8013 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
abstract
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.
Dianbing Xi, Jiepeng Wang 0005, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang 0004, Chi Zhang 0012, Xuelong Li 0001
AAAI3
2025 InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild
Yiyi Ma, Yuanzhi Liang, Xiu Li 0001, Chi Zhang 0012, Xuelong Li 0001
ICCV2
2025 Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
abstract
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene—within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
Sheng Liu 0013, Yuanzhi Liang, Jiepeng Wang 0005, Sidan Du, Chi Zhang 0067, Xuelong Li 0001
SIGGRAPH Asia2
2024 FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
abstract
Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing long video diffusion models. This paper investigates a straightforward and training-free approach to extend an existing short video diffusion model (e.g. pre-trained on 16-frame videos) for consistent long video generation (e.g. 128 frames). Our preliminary observation has found that directly applying the short video diffusion model to generate long videos can lead to severe video quality degradation. Further investigation reveals that this degradation is primarily due to the distortion of high-frequency components in long videos, characterized by a decrease in spatial high-frequency components and an increase in temporal high-frequency components. Motivated by this, we propose a novel solution named FreeLong to balance the frequency distribution of long video features during the denoising process. FreeLong blends the low-frequency components of global video features, which encapsulate the entire video sequence, with the high-frequency components of local video features that focus on shorter subsequences of frames. This approach maintains global consistency while incorporating diverse and high-quality spatiotemporal details from local videos, enhancing both the consistency and fidelity of long video generation. We evaluated FreeLong on multiple base video diffusion models and observed significant improvements. Additionally, our method supports coherent multi-prompt generation, ensuring both visual coherence and seamless transitions between scenes. Our project page is at: https://yulu.net.cn/freelong.
Yu Lu 0019, Yuanzhi Liang, Linchao Zhu, Yi Yang 0001
NeurIPS2
2024 Split-Check: Boosting Product Recognition via Instance-Level Retrieval
abstract
AI-based methods are shining across a variety of industries, especially unmanned retail. Product recognition is the problem of recognizing the category and quantity of products (e.g., beverages and mineral water) in intelligent unmanned vending machines (UVMs) to automatic checkout during purchase. However, for similar products in hundreds of categories, the existing method is not accurate enough. Besides, they cannot be extended for new products without retraining. In this article, we propose a product recognition approach based on intelligent UVMs, calledSplit-Check, which first splits the region of interest of products by detection and then check product by instance-level retrieval. Split-Check is the combination of two important components. The preliminary detection distinguishes items that contain the different coarse-grained features, then locates items, and classifies them into coarse-grained categories as a candidate. The retrieval further distinguishes the candidate items that contain the different fine-grained features. Besides, we reconstruct a large-scale categories product dataset GOODS-85 based on actual UVMs scenarios, in which the number of categories of items is larger than the existing dataset. Experimental results demonstrate the effectiveness of the proposed approach. Our method significantly improves the recognition performance of hundreds of products and increases the scalability of products.
Chengxu Liu 0001, Zongyang Da, Yuanzhi Liang, Guoshuai Zhao 0001, Xueming Qian
IEEE Trans. Ind. Informatics3
2024 IcoCap: Improving Video Captioning by Compounding Images
abstract
Video captioning is a more challenging task compared to image captioning, primarily due to differences in content density. Video data contains redundant visual content, making it difficult for captioners to generalize diverse content and avoid being misled by irrelevant elements. Moreover, redundant content is not well-trimmed to match the corresponding visual semantics in the ground truth, further increasing the difficulty of video captioning. Current research in video captioning predominantly focuses on captioner design, neglecting the impact of content density on captioner performance. Considering the differences between videos and images, there exists an another line to improve video captioning by leveraging concise and easily-learned image samples to further diversify video samples. This modification to content density compels the captioner to learn more effectively against redundancy and ambiguity. In this article, we propose a novel approach calledImage-Compounded learning for videoCaptioners (IcoCap) to facilitate better learning of complex video semantics. IcoCap comprises two components: the Image-Video Compounding Strategy (ICS) and Visual-Semantic Guided Captioning (VGC). ICS compounds easily-learned image semantics into video semantics, further diversifying video content and prompting the network to generalize contents in a more diverse sample. Besides, learning with the sample compounded with image contents, the captioner is compelled to better extract valuable video cues in the presence of straightforward image semantics. This helps the captioner further focus on relevant information while filtering out extraneous content. Then, VGC guides the network in flexibly learning ground truth captions based on the compounded samples, helping to mitigate the mismatch between the ground truth and ambiguous semantics in video samples. Our experimental results demonstrate the effectiveness of IcoCap in improving the learning of video captioners. Applied to the widely-used MSVD, MSR-VTT, and VATEX datasets, our approach achieves competitive or superior results compared to state-of-the-art methods, illustrating its capacity to handle the redundant and ambiguous video data
Yuanzhi Liang, Linchao Zhu, Yi Yang 0001
IEEE Trans. Multim.1
2024 Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification
abstract
Though significant progress has been achieved on fine-grained visual classification (FGVC), severe overfitting still hinders model generalization. A recent study shows that hard samples in the training set can be easily fit, but most existing FGVC methods fail to classify some hard examples in the test set. The reason is that the model overfits those hard examples in the training set, but does not learn to generalize to unseen examples in the test set. In this article, we propose a moderate hard example modulation (MHEM) strategy to properly modulate the hard examples. MHEM encourages the model to not overfit hard examples and offers better generalization and discrimination. First, we introduce three conditions and formulate a general form of a modulated loss function. Second, we instantiate the loss function and provide a strong baseline for FGVC, where the performance of a naive backbone can be boosted and be comparable with recent methods. Moreover, we demonstrate that our baseline can be readily incorporated into the existing methods and empower these methods to be more discriminative. Equipped with our strong baseline, we achieve consistent improvements on three typical FGVC datasets, i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft. We hope the idea of moderate hard example modulation will inspire future research work toward more effective fine-grained visual recognition.
Yuanzhi Liang, Linchao Zhu, Yi Yang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Product Recognition for Unmanned Vending Machines
abstract
Recently, the emerging concept of "unmanned retail" has drawn more and more attention, and the unmanned retail based on the intelligent unmanned vending machines (UVMs) scene has great market demand. However, existing product recognition methods for intelligent UVMs cannot adapt to large-scale categories and have insufficient accuracy. In this article, we propose a method for large-scale categories product recognition based on intelligent UVMs. It can be divided into two parts: 1) first, we explore the similarities and differences between products through manifold learning, and then we build a hierarchical multigranularity label to constrain the learning of representation; and 2) second, we propose a hierarchical label object detection network, which mainly includes coarse-to-fine refine module (C2FRM) and multiple granularity hierarchical loss (MGHL), which are used to assist in capturing multigranularity features. The highlights of our method are mine potential similarity between large-scale category products and optimization through hierarchical multigranularity labels. Besides, we collected a large-scale product recognition dataset GOODS-85 based on the actual UVMs scenario. Experimental results and analysis demonstrate the effectiveness of the proposed product recognition methods.
Chengxu Liu 0001, Zongyang Da, Yuanzhi Liang, Guoshuai Zhao 0001, Xueming Qian
IEEE Trans. Neural Networks Learn. Syst.3
2023 MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects
abstract
Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to act and how to act. Correspondingly, the task mainly requires producing actionability scores, action proposals, and success likelihood scores according to the given 3D object information and robotic information. Current works usually directly process multi-modal inputs with early fusion and apply critic networks to produce scores, which leads to insufficient multi-modal learning ability and inefficiently iterative training in multiple stages. This paper proposes a novel Multimodality-Aware Autoencoder-based affordance Learning (MAAL) for the 3D object affordance problem. It is an efficient pipeline, trained in one go, and only requires a few positive samples in training data. More importantly, MAAL contains a MultiModal Energized Encoder (MME) for better multi-modal learning. It comprehensively models all multi-modal inputs from 3D objects and robotic actions. Jointly considering information from multiple modalities, the encoder further learns interactions between robots and objects. MME empowers the better multi-modal learning ability for understanding object affordance. Experimental results and visualizations, based on a large-scale dataset PartNet-Mobility, show the effectiveness of MAAL in learning multi-modal data and solving the 3D articulated object affordance problem.
Yuanzhi Liang, Linchao Zhu, Yi Yang 0001
ICCV1
2023 Anomaly detection framework for unmanned vending machines
Zongyang Da, Yujie Dun, Chengxu Liu 0001, Yuanzhi Liang, Xueming Qian
Knowl. Based Syst.4
2022 A Knowledge-Guided Method for Disease Prediction Based on Attention Mechanism
Yuanzhi Liang, Haofen Wang
WISA1
2022 SEEG: Semantic Energized Co-speech Gesture Generation
abstract
Talking gesture generation is a practical yet challenging task that aims to synthesize gestures in line with speech. Gestures with meaningful signs can better convey useful information and arouse sympathy in the audience. Current works focus on aligning gestures with the speech rhythms, which are difficult to mine the semantics and model semantic gestures explicitly. This paper proposes a novel semantic Energized Generation (SEEG) method for semantic-aware gesture generation. Our method contains two parts: DEcoupled Mining module (DEM) and Semantic Energizing Module (SEM). DEM decouples the semantic-irrelevant information from inputs and separately mines information for the beat and semantic gestures. SEM conducts semantic learning and produces semantic gestures. Apart from representational similarity, SEM requires the predictions to express the same semantics as the ground truth. Besides, a semantic prompter is designed in SEM to leverage the semantic-aware supervision to predictions. This promotes the networks to learn and generate semantic gestures. Experimental results reported in three metrics on different benchmarks prove that SEEG efficiently mines semantic cues and generates semantic gestures. SEEG outperforms other methods in all semantic-aware evaluations on different datasets. Qualitative evaluations also indicate the superiority of SEEG in semantic expressiveness. Code is available via https://github.com/akira-l/SEEG.
Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Yi Yang 0001
CVPR1
2022 A Simple Episodic Linear Probe Improves Visual Recognition in the Wild
abstract
Understanding network generalization and feature discrimination is an open research problem in visual recognition. Many studies have been conducted to assess the quality of feature representations. One of the simple strategies is to utilize a linear probing classifier to quantitatively evaluate the class accuracy under the obtained features. The typical linear probe is only applied as a proxy at the inference time, but its efficacy in measuring features' suitability for linear classification is largely neglected in training. In this paper, we propose an episodic linear probing (ELP) classifier to reflect the generalization of visual rep-resentations in an online manner. ELP is trained with detached features from the network and re-initialized episodically. It demonstrates the discriminability of the visual representations in training. Then, an ELP-suitable Regularization term (ELP-SR) is introduced to reflect the distances of probability distributions between the ELP classifier and the main classifier. ELP-SR leverages are-scaling factor to regularize each sample in training, which modulates the loss function adaptively and encourages the features to be discriminative and generalized. We observe significant improvements in three real-world visual recognition tasks: fine-grained visual classification, long-tailed visual recognition, and generic object recognition. The performance gains show the effectiveness of our method in im-proving network generalization and feature discrimination.
Yuanzhi Liang, Linchao Zhu, Yi Yang 0001
CVPR1
2022 Shyper: An embedded hypervisor applying hierarchical resource isolation strategies for mixed-criticality systems
abstract
With the development of the IoT, modern embedded systems are evolving to general-purpose and mixed-criticality systems, where virtualization has become the key to guarantee the isolation between tasks with different criticality. Traditional server-based hypervisors (KVM and Xen) are difficult to use in embedded scenarios due to performance and security reasons. As a result, several new hypervisors (Jailhouse and Bao) have been proposed in recent years, which effectively solve the problems above through static partitioning. However, this inflexible resource isolation strategy assumes no resource sharing across guests, which greatly reduces the resource utilization and VM scalability. This prevents themselves from simultaneously fulfilling the differentiated demands from VMs conducting different tasks. This paper proposes an efficient and real-time embedded hypervisor “Shyper”, aiming at providing differentiated services for VMs with different criticality. To achieve that, Shyper supports fine-grained hierarchical resource isolation strategies and introduces several novel “VM-Exit-less” real-time virtualization techniques, which grants users the flexibility to strike a trade-off between VM's resource utilization and real-time performance. In this paper, we also compare Shyper with other mainstream hypervisors (KVM, Jailhouse, etc.) to evaluate its feasibility and effectiveness.
Yicong Shen, Lei Wang 0126, Yuanzhi Liang, Bo Jiang 0001
DATE3
2022 Deep Knowledge Reasoning guided Disease Prediction
abstract
Disease prediction, which aims to predict possible future diseases for patients, is a fundamental research problem in medical informatics. Many studies have proposed the introduction of external knowledge to enhance deep learning models which have achieved good results, but as most of these studies only consider using single-hop relationship information from knowledge graphs or simply introduce partial knowledge graphs, fail to effectively mine knowledge paths to understand the relationships between diseases. To this end, we propose a new approach, which uses existing medical knowledge graphs for multi-hop reasoning to guide the self-attention based transformer model for disease prediction. Specifically, we design a reinforcement learning algorithm to perform path reasoning in the knowledge graph to obtain explicit disease progression paths and fusion them with the original Electronic Health Records (EHR) data. After embedding we capture the implicit relationships between the deep knowledge and the original EHR information through several transformer encoders based on the self-attention mechanism to better extract features. At the same time, multi-hop knowledge is more interpretable than single-hop knowledge in terms of disease prediction. Experimental results on the real-world medical dataset MIMIC-III show the superiority of the proposed approach compared to a series of state-of-the-art baselines.11The code has been uploaded to https://github.com/15536385781/DKRDP
Yuanzhi Liang, Haofen Wang
SMC1
2021 Removing Raindrops and Rain Streaks in One Go
abstract
Existing rain-removal algorithms often tackle either rain streak removal or raindrop removal, and thus may fail to handle real-world rainy scenes. Besides, the lack of real-world deraining datasets comprising different types of rain and their corresponding rain-free ground-truth also impedes deraining algorithm development. In this paper, we aim to address real-world deraining problems from two aspects. First, we propose a complementary cascaded network architecture, namely CCN, to remove rain streaks and raindrops in a unified framework. Specifically, our CCN removes raindrops and rain streaks in a complementary fashion, i.e., raindrop removal followed by rain streak removal and vice versa, and then fuses the results via an attention based fusion module. Considering significant shape and structure differences between rain streaks and raindrops, it is difficult to manually design a sophisticated network to remove them effectively. Thus, we employ neural architecture search to adaptively find optimal architectures within our specified deraining search space. Second, we present a new real-world rain dataset, namely RainDS, to prosper the development of deraining algorithms in practical scenarios. RainDS consists of rain images in different types and their corresponding rain-free ground-truth, including rain streak only, raindrop only, and both of them. Extensive experimental results on both existing benchmarks and RainDS demonstrate that our method outperforms the state-of-the-art.
Ruijie Quan, Xin Yu 0002, Yuanzhi Liang, Yi Yang 0001
CVPR3
2021 Towards Better Railway Service: Passengers Counting in Railway Compartment
abstract
Counting passengers in railway compartments is an essential problem for improving service quality, user experience, public security, and disaster relief in the railway system. Considering many limitations in the compartment, the infrared sensor, 3D camera, etc. are not practical in this scene. Due to the flexibility and lower cost, solutions with standard cameras attract much attention in real applications. However, since the problem with scale variation in the narrow space is different from universal detection or counting problems, the specific benchmark of dataset and methods should be provided and proposed for this task. In this paper, we provide a passenger counting dataset. Relying on this dataset, we propose a passenger counting method. The solution contains a motion supervised multi-scale representation method which provides proposals against scale variation, a spatially-temporally enhanced counting which provides precise counting numbers, and a partial proposal method which conducts methods to be utilized in reality. With the proposed solution, the passengers counting task is solved in higher accuracy and practicable in the compartment environment. In experiments, the results show that all the modules in our solution are useful and efficient, and our method outperforms in comparison with others in the compartment scene.
Yuanzhi Liang, Xueming Qian, Li Zhu 0003
IEEE Trans. Circuits Syst. Video Technol.1
2021 Food and Ingredient Joint Learning for Fine-Grained Recognition
abstract
Fine-grained food recognition is the detailed classification that provides more specialized and professional attribute information of food. It is the basic work to realize healthy diet recommendations and cooking instructions, nutrition intake management, and cafeteria self-checkout system. Chinese food lacks structured information, and ingredients composition is an important consideration. The current approaches mostly focus on global dish appearance without any analysis of ingredient composition and fully considering the attention of regional features. In this paper, we propose an Attention Fusion Network (AFN) and Food-Ingredient Joint Learning module for fine-grained food and ingredients recognition. The AFN first focuses on the food discrimination region against unstructured defeat and generates the feature embeddings jointly aware of the ingredients and food. The Food-Ingredient Joint Learning module aims at alleviating the issue of ingredients imbalance. Therefore, we propose a balance focal loss to optimize the feature expression ability of the network for ingredients. In experiments, the results of ingredients recognition show the state-of-the-art performances on fine-grained Chinese food dataset VIREO Food-172.
Chengxu Liu 0001, Yuanzhi Liang, Xueming Qian, Jianlong Fu
IEEE Trans. Circuits Syst. Video Technol.2
2019 VrR-VG: Refocusing Visually-Relevant Relationships
abstract
Relationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than ``learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon.
Yuanzhi Liang, Yalong Bai, Wei Zhang 0031, Xueming Qian, Li Zhu 0003, Tao Mei 0001
ICCV1
2016 A self-adapting method for RBC count from different blood smears based on PCNN and image quality
abstract
Microscopic image processing is critical aspects to biomedical image analysis, and blood cell counts are very important role in medical diagnoses. Various dyeing methods and microscopes are used, so we need methods that can effectively count cells by adapting to such diversity. This paper presents a new method that extracts the contours of red blood cells based on the quality of a binary image that is preprocessed using PCNN. The method solves the various blood smear issues caused by the different cell dyeing methods. Moreover, it uses a self-adapting method for counting cells, using the circular Hough transform(CHT) for different amplifications. The experimental results show that the proposed method performed better in contrast variations between cells and background. The method is also much more efficient in segmentation on overlapped cells, and much more accurate in counting RBC results.
Yuanzhi Liang, Yide Ma
BIBM2