Ziyun Qian

dblp:339/1300 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0007-9800-1253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
abstract
With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods rely on pre-trained text-to-image diffusion models and monocular depth estimators. However, the generated scenes occupy large amounts of storage space and often lack effective regularisation methods, leading to geometric distortions. To this end, we propose BloomScene, a lightweight structured 3D Gaussian splatting for crossmodal scene generation, which creates diverse and high-quality 3D scenes from text or image inputs. Specifically, a crossmodal progressive scene generation framework is proposed to generate coherent scenes utilizing incremental point cloud reconstruction and 3D Gaussian splatting. Additionally, we propose a hierarchical depth prior-based regularization mechanism that utilizes multi-level constraints on depth accuracy and smoothness to enhance the realism and continuity of the generated scenes. Ultimately, we propose a structured context-guided compression mechanism that exploits structured hash grids to model the context of unorganized anchor attributes, which significantly eliminates structural redundancy and reduces storage overhead. Comprehensive experiments across multiple scenes demonstrate the significant potential and advantages of our framework compared with several baselines.
Xiaolu Hou, Mingcheng Li, Dingkang Yang, Jiawei Chen 0012, Ziyun Qian, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
AAAI5
2025 MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
abstract
Diffusion models have shown excellent performance in text-to-image generation. Nevertheless, existing methods often suffer from performance bottlenecks when handling complex prompts that involve multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Compositional Diffusion (MCCD) for text-to-image generation for complex scenes. Specifically, we design a multi-agent collaboration-based scene parsing module that generates an agent system comprising multiple agents with distinct tasks, utilizing MLLMs to extract various scene elements effectively. In addition, Hierarchical Compositional diffusion utilizes a Gaussian mask and filtering to refine bounding box regions and enhance objects through region enhancement, resulting in the accurate and high-fidelity generation of complex scenes. Comprehensive experiments demonstrate that our MCCD significantly improves the performance of the baseline models in a training-free manner, providing a substantial advantage in complex scene generation.
Mingcheng Li, Xiaolu Hou, Dingkang Yang, Ziyun Qian, Jiawei Chen 0012, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
CVPR5
2025 MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion
abstract
Motion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine details in complex motions, leading to models that capture only coarse-grained style characteristics. To overcome this limitation, we introduce the Motion Adaptive Fusion Diffusion (MAFD) framework, which leverages adaptive signal fusion to highlight essential style-defining features while minimizing redundant information. Moreover, current diffusion-based denoisers often fail to effectively capture the temporal relationships in motion sequences, producing rigid and fragmented stylized motions. Drawing inspiration from the Mamba model, we propose the Style Mamba Denoiser (SMD), which adopts a selection mechanism to preserve long-range dependencies and maintain temporal coherence. Extensive experiments show that our approach outperforms state-of-the-art methods in both qualitative and quantitative evaluations, achieving more refined and coherent stylized motions.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Dongliang Kou, Lihua Zhang 0002
ICASSP1
2025 VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
abstract
In histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc. ) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution for generating virtually stained images, significantly reducing the time and labor costs associated with traditional histochemical staining. However, a new challenge arises in separating the fundamental visual characteristics of tissue sections from the visual differences induced by staining agents. Additionally, virtual staining often overlooks essential pathological knowledge and the physical properties of staining, resulting in only style-level transfer. To address these issues, we introduce, for the first time in virtual staining tasks, a pathological vision-language large model (VLM) as an auxiliary tool. We integrate contrastive learnable prompts, foundational concept anchors for tissue sections, and staining-specific concept anchors to leverage the extensive knowledge of the pathological VLM. This approach is designed to describe, frame, and enhance the direction of virtual staining. Furthermore, we have developed a data augmentation method based on the constraints of the VLM. This method utilizes the VLM's powerful image interpretation capabilities to further integrate image style and structural information, proving beneficial in high-precision pathological diagnostics. Extensive evaluations on publicly available multi-domain unpaired staining datasets demonstrate that our method can generate highly realistic images and enhance the accuracy of downstream tasks, such as glomerular detection and segmentation. Our code. https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR is available.
Zizhi Chen, Minghao Han, Yizhou Liu 0002, Ziyun Qian, Xukun Zhang, Jingwei Wei, Lihua Zhang 0002
ACM Multimedia5
2025 UMSD: High Realism Motion Style Transfer via Unified Mamba-based Diffusion
abstract
Motion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer.
Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002
ACM Multimedia1
2025 MSTDF: Motion Style Transfer Towards High Visual Fidelity Based on Dynamic Fusion
abstract
Emotion-guided motion style transfer is a novel research direction, enabling the efficient generation of motion in various emotional styles for use in films, games, and other domains. However, existing methods primarily rely on global feature statistics for motion style transfer, neglecting local semantic structure and resulting in the degradation of motion content structure. This letter proposes a novel Motion Style Transfer based on Dynamic Fusion (MSTDF) framework, which treats content and style motion as distinct signals and employs dynamic fusion for high-fidelity motion style transfer. Additionally, to address the challenge of traditional discriminators capturing subtle motion style features, we propose the Motion Dynamic Fusion (MDF) discriminator to capture the details and fine-grained style characteristics of motion sequences, assisting the generator in producing higher-fidelity stylized motion. Finally, extensive experiments on the Xia dataset demonstrate that our method surpasses state-of-the-art methods in qualitative and quantitative comparisons.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Zeyu Xiao 0001, Lihua Zhang 0002
IEEE Signal Process. Lett.1
2024 SceneWeaver: Text-Driven Scene Generation with Geometry-aware Gaussian Splatting
Xiaolu Hou, Mingcheng Li, Jiawei Chen 0012, Dingkang Yang, Ziyun Qian, Lihua Zhang 0002
ACML5
2024 Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
abstract
Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines.
Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002
CVPR9
2024 Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, Lihua Zhang 0002
MICCAI (5)6
2024 Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
abstract
Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases.
Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002
NeurIPS12
2023 HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
abstract
We propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose information, which contains hand prior knowledge that is strongly correlated with occluded regions. Thus, in this work, we propose a novel 3D hand mesh reconstruction network HandGCAT, that can fully exploit hand prior as compensation information to enhance occluded region features. Specifically, we designed the Knowledge-Guided Graph Convolution (KGC) module and the Cross-Attention Transformer (CAT) module. KGC extracts hand prior information from 2D hand pose by graph convolution. CAT fuses hand prior into occluded regions by considering their high correlation. Extensive experiments on popular datasets with challenging hand-object occlusions, such as HO3D v2, HO3D v3, and DexYCB demonstrate that our HandGCAT reaches state-of-the-art performance. The code is available at https://github.com/heartStrive/HandGCAT.
Shuaibing Wang, Shunli Wang 0001, Dingkang Yang, Mingcheng Li, Ziyun Qian, Liuzhen Su, Lihua Zhang 0002
ICME5