Muyang Zhang

dblp:18/4368 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Robust detection in complex construction sites: HiPA-DETR with weather-aware and cross-domain generalization
Zenghuang Fu, Muyang Zhang, Changwei Wang 0001, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.5
2025 PanoDiT: Panoramic Videos Generation with Diffusion Transformer
abstract
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose PanoDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNet-based denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 0001, Weiliang Meng, Jianwei Guo 0003, Xiaopeng Zhang 0001
AAAI1
2025 HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
abstract
Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from existing videos, thereby lacking flexibility. To address this, we propose HumanDreamer, a decoupled human video generation framework that first generates diverse poses from text prompts and then leverages these poses to generate human-motion videos. Specifically, we propose MotionVid, the largest dataset for human-motion pose generation. Based on the dataset, we present MotionDiT, which is trained to generate structured human-motion poses from text prompts. Besides, a novel LAMA loss is introduced, which together contribute to a significant improvement in FID by 62.4%, along with respective enhancements in R-precision for top1, top2, and top3 by 41.8%, 26.3%, and 18.3%, thereby advancing both the Text-to-Pose control accuracy and FID metrics. Our experiments across various Pose-to-Video baselines demonstrate that the poses generated by our method can produce diverse and high-quality human-motion videos. Furthermore, our model can facilitate other downstream tasks, such as pose sequence prediction and 2D-3D motion lifting.
Chaojun Ni, Guosheng Zhao, Zhiqin Yang, Muyang Zhang, Xinze Chen, Guan Huang 0003, Lihong Liu, Xingang Wang 0003
CVPR7
2025 AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD
Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao
DAC3
2025 AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention
abstract
With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in this field. To address these limitations, we introduce AccidentX, a large-scale multimodal dataset specifically curated for comprehensive traffic accident analysis and prevention. Our AccidentX comprises over 10,000 bird’s-eye view (BEV) videos generated using the CARLA simulator, with detailed annotations covering a wide range of traffic scenarios. In comparison to existing datasets such as nuScenes, our AccidentX offers seven times more video frames and leverages Vision-Language Models (VLMs) and GPT-4o for enhanced scene understanding and decision-making. We also establish a benchmark for state-of-the-art Multimodal Large Language Models (MLLMs) on AccidentX, fostering further research and innovation within the community. AccidentX will be made available as a fully open source resource for the advancement of the autonomous driving safety algorithm community.
Muyang Zhang, Mingda Jia, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
IROS1
2025 EventVAD: Training-Free Event-Aware Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs to perform fine-grained temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning and making final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs. The code is available at https://github.com/YihuaJerry/EventVAD.
Yihua Shao, Haojin He, Siyu Chen 0021, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma 0005, Hao Tang 0005, Yan Wang 0105, Shuyan Li
ACM Multimedia8
2025 PDFT: parameter-diminish fine-tuning for transformer-based models
Muyang Zhang, Weiliang Meng, Mingda Jia, Jiaming Gu, Yihua Shao, Changwei Wang 0001, Rongtao Xu, Xiaopeng Zhang 0001
Vis. Comput.1
2024 AG-SDM: Aquascape generation based on stable diffusion model with low-rank adaptation
abstract
Abstract As an amalgamation of landscape design and ichthyology, aquascape endeavors to create visually captivating aquatic environments imbued with artistic allure. Traditional methodologies in aquascape, governed by rigid principles such as composition and color coordination, may inadvertently curtail the aesthetic potential of the landscapes. In this paper, we propose Aquascape Generation based on Stable Diffusion Models (AG‐SDM), prioritizing aesthetic principles and color coordination to offer guiding principles for real artists in Aquascape creation. We meticulously curated and annotated three aquascape datasets with varying aspect ratios to accommodate diverse landscape design requirements regarding dimensions and proportions. Leveraging the Fréchet Inception Distance (FID) metric, we trained AGFID for quality assessment. Extensive experiments validate that our AG‐SDM excels in generating hyper‐realistic underwater landscape images, closely resembling real flora, and achieves state‐of‐the‐art performance in aquascape image generation.
Muyang Zhang, Yuewei Xian, Wei Li 0237, Jiaming Gu, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds1
2023 FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose Conditions
abstract
Collaborative perception offers a promising solution to overcome challenges such as occlusion and long-range data processing. However, limited sensor accuracy leads to noisy poses that misalign observations among vehicles. To address this problem, we propose the FeaCo, which achieves robust Feature-level Consensus among collaborating agents in noisy pose conditions without additional training. We design an efficient Pose-error Rectification Module (PRM) to align derived feature maps from different vehicles, reducing the adverse effect of noisy pose and bandwidth requirements. We also provide an effective multi-scale Cross-level Attention Module (CAM) to enhance information aggregation and interaction between various scales. Our FeaCo outperforms all other localization rectification methods, as validated on both the collaborative perception simulation dataset OPV2V and real-world dataset V2V4Real, reducing heading error and enhancing localization accuracy across various error levels. Our code is available at: https://github.com/jmgu0212/FeaCo.git.
Jiaming Gu, Muyang Zhang, Weiliang Meng, Shibiao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
ACM Multimedia3
2023 HTCViT: an effective network for image classification and segmentation based on natural disaster datasets
Wei Li 0237, Muyang Zhang, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.3
2013 A Divide-and-Conquer Approach to Quad Remeshing
abstract
Many natural and man-made objects consist of simple primitives, similar components, and various symmetry structures. This paper presents a divide-and-conquer quadrangulation approach that exploits such global structural information. Given a model represented in triangular mesh, we first segment it into a set of submeshes, and compare them with some predefined quad mesh templates. For the submeshes that are similar to a predefined template, we remesh them as the template up to a number of subdivisions. For the others, we adopt the wave-based quadrangulation technique to remesh them with extensions to preserve symmetric structure and generate compatible quad mesh boundary. To ensure that the individually remeshed submeshes can be seamlessly stitched together, we formulate a mixed-integer optimization problem and design a heuristic solver to optimize the subdivision numbers and the size fields on the submesh boundaries. With this divider-and-conquer quadrangulation framework, we are able to process very large models that are very difficult for the previous techniques. Since the submeshes can be remeshed individually in any order, the remeshing procedure can run in parallel. Experimental results showed that the proposed method can preserve the high-level structures, and process large complex surfaces robustly and efficiently.
Muyang Zhang, Jin Huang 0001, Xinguo Liu, Hujun Bao
IEEE Trans. Vis. Comput. Graph.1
2011 Controllable highly regular triangulation
Jin Huang 0001, Muyang Zhang, Wenjie Pei, Wei Hua 0002, Hujun Bao
Sci. China Inf. Sci.2
2010 A wave-based anisotropic quadrangulation method
abstract
This paper proposes a new method for remeshing a surface into anisotropically sized quads. The basic idea is to construct a special standing wave on the surface to generate the global quadrilateral structure. This wave based quadrangulation method is capable of controlling the quad size in two directions and precisely aligning the quads with feature lines. Similar to the previous methods, we augment the input surface with a vector field to guide the quad orientation. The anisotropic size control is achieved by using two size fields on the surface. In order to reduce singularity points, the size fields are optimized by a new curl minimization method. The experimental results show that the proposed method can successfully handle various quadrangulation requirements and complex shapes, which is difficult for the existing state-of-the-art methods.
Muyang Zhang, Jin Huang 0001, Xinguo Liu, Hujun Bao
ACM Trans. Graph.1
2008 Spectral quadrangulation with orientation and alignment control
abstract
This paper presents a new quadrangulation algorithm, extending the spectral surface quadrangulation approach where the coarse quadrangular structure is derived from the Morse-Smale complex of an eigenfunction of the Laplacian operator on the input mesh. In contrast to the original scheme, we provide flexible explicit controls of the shape, size, orientation and feature alignment of the quadrangular faces. We achieve this by proper selection of the optimal eigenvalue (shape), by adaption of the area term in the Laplacian operator (size), and by adding special constraints to the Laplace eigenproblem (orientation and alignment). By solving a generalized eigen-problem we can generate a scalar field on the mesh whose Morse-Smale complex is of high quality and satisfies all the user requirements. The final quadrilateral mesh is generated from the Morse-Smale complex by computing a globally smooth parametrization. Here we additionally introduce edge constraints to preserve user specified feature lines accurately.
Jin Huang 0001, Muyang Zhang, Xinguo Liu, Leif Kobbelt, Hujun Bao
ACM Trans. Graph.2