Lingzhou Mu

dblp:331/1421 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0001-0184-0664ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 44% Multi-agent systems · 32% 3D vision · 20%
Network and information security
1 paper
Security and privacy of machine learning · 92% Digital forensics and information hiding · 8%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › human body modeling › 3d human modeling
human-scene interaction synthesis
1.012026
FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
LLM-based multi-agent planning
1.012026
FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
1.012026
FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework · AAAI 2026
Machine learning › Generative modeling
video generation
1.012026
FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework · AAAI 2026
Machine learning › Generative modeling › diffusion model › controllable generation
controllable diffusion generation
0.912025
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model · ACM Multimedia 2025
Machine learning › Generative modeling
diffusion model
0.912025
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model · ACM Multimedia 2025
Visual content generation and editing › talking head generation
audio-driven portrait animation
0.912025
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model · ACM Multimedia 2025
Visual content generation and editing
talking head generation
0.912025
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model · ACM Multimedia 2025
Security and privacy of machine learning
adversarial attack
0.612022
Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations? · ACM Multimedia 2022
Security and privacy of machine learning › model intellectual property protection
model watermarking
0.612022
Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations? · ACM Multimedia 2022
Security and privacy of machine learning
watermark removal attack
0.612022
Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations? · ACM Multimedia 2022
Robotics › Motion planning and robot control
path planning
0.312026
FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework · AAAI 2026
Computer vision › 3D vision › 3d shape modeling
3d head modeling
0.312025
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model · ACM Multimedia 2025
Security and privacy of machine learning › model intellectual property protection › model watermarking
deep neural network watermarking
0.212022
Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations? · ACM Multimedia 2022
Digital forensics and information hiding
watermarking
0.212022
Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations? · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.73d head parameter conditioning · 1.7large language model · 1.0graph-based multi-agent framework · 1.0direct preference optimization · 1.0input preprocessing · 0.6adversarial perturbation · 0.6
YearPublicationVenuePosition
2026 FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework
abstract
Human-Scene Interaction (HSI) seeks to generate realistic human behaviors within complex environments, yet it faces significant challenges in handling long-horizon, high-level tasks and generalizing to unseen scenes. To address these limitations, we introduce FantasyHSI, a novel HSI framework centered on video generation and multi-agent systems that operates without paired data. We model the complex interaction process as a dynamic directed graph, upon which we build a collaborative multi-agent system. This system comprises a scene navigator agent for environmental perception and high-level path planning, and a planning agent that decomposes long-horizon goals into atomic actions. Critically, we introduce a critic agent that establishes a closed-loop feedback mechanism by evaluating the deviation between generated actions and the planned path. This allows for the dynamic correction of trajectory drifts caused by the stochasticity of the generative model, thereby ensuring long-term logical consistency. To enhance the physical realism of the generated motions, we leverage Direct Preference Optimization (DPO) to train the action generator, significantly reducing artifacts such as limb distortion and foot-sliding. Extensive experiments on our custom SceneBench benchmark demonstrate that FantasyHSI significantly outperforms existing methods in terms of generalization, long-horizon task completion, and physical realism.
Lingzhou Mu, Mengchao Wang, Mu Xu
AAAI1
2025 FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
abstract
Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor controllability, making them less applicable to real-world scenarios such as filmmaking and live streaming for e-commerce. To address this limitation, we propose FLAP, a novel approach that integrates explicit 3D intermediate parameters (head poses and facial expressions) into the diffusion model for end-to-end generation of realistic portrait videos. The proposed architecture allows the model to generate vivid portrait videos from audio while simultaneously incorporating additional control signals, such as head rotation angles and eye-blinking frequency. Furthermore, the decoupling of head pose and facial expression allows for independent control of each, offering precise manipulation of both the avatar's pose and facial expressions. We also demonstrate its flexibility in integrating with existing 3D head generation methods, bridging the gap between 3D model-based approaches and end-to-end diffusion techniques. Extensive experiments show that our method outperforms recent audio-driven portrait video models in both naturalness and controllability.
Lingzhou Mu, Baiji Liu, Guiming Mo, Jiawei Jin, Kai Zhang 0012, Hao-Zhi Huang 0001
ACM Multimedia1
2022 Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?
abstract
Training Deep Neural Networks (DNN) is a time-consuming process and requires a large amount of training data, which motivates studies working on protecting the intellectual property (IP) of DNN models by employing various watermarking techniques. Unfortunately, in recent years, adversaries have been exploiting the vulnerabilities of the employed watermarking techniques to remove the embedded watermarks. In this paper, we investigate and introduce a novel watermark removal attack, called AdvNP, against all the existing four different types of DNN watermarking schemes via input preprocessing by injecting Adversarial Naturalness-aware Perturbations. In contrast to the prior studies, our proposed method is the first work that generalizes all the existing four watermarking schemes well without involving any model modification, which preserves the fidelity of the target model. We conduct the experiments against four state-of-the-art (SOTA) watermarking schemes on two real tasks (e.g., image classification on ImageNet, face recognition on CelebA) across multiple DNN models. Overall, our proposed AdvNP significantly invalidates the watermarks against the four watermarking schemes on two real-world datasets, i.e., 60.9% on the average attack success rate and up to 97% in the worse case. Moreover, our AdvNP could well survive the image denoising techniques and outperforms the baseline in both the fidelity preserving and watermark removal. Furthermore, we introduce two defense methods to enhance the robustness of DNN watermarking against our AdvNP. Our experimental results pose real threats to the existing watermarking schemes and call for more practical and robust watermarking techniques to protect the copyright of pre-trained DNN models. The source code and models are available at ttps://github.com/GitKJ123/AdvNP.
Run Wang 0001, Lingzhou Mu, Jixing Ren, Shangwei Guo, Liming Fang 0001, Jing Chen 0003, Lina Wang 0001
ACM Multimedia3