Hengyu Meng

dblp:334/0298 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0005-8961-3982ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 56% Geometric modeling and processing · 28% Virtual and augmented reality · 17%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
3d content creation
0.912025
Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D Sculpting · ICCV 2025
Visual content generation and editing
image generation
0.912025
MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image Generation · VR 2025
Virtual and augmented reality
immersive interaction
0.312025
MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image Generation · VR 2025
Virtual and augmented reality › interactive storytelling
immersive storytelling
0.312025
MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image Generation · VR 2025

Methods — techniques the papers use, named apart from their topics

style control · 0.9semantic-aware denoising · 0.9score distillation sampling · 0.9prompt token blending · 0.9layout prediction · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2025 DEGAS: Detailed Expressions on Full-Body Gaussian Avatars
abstract
Although neural rendering has made significant ad-vances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS, the first 3D Gaussian Splatting (3DGS)-based modeling method for full-body avatars with rich facial expressions. Trained on multiview videos of a given subject, our method learns a conditional variational autoencoder that takes both the body motion and facial expression as driving signals to generate Gaussian maps in the UV layout. To drive the facial expressions, instead of the commonly used 3D Mor-phable Models (3DMMs) in 3D head avatars, we propose to adopt the expression latent space trained solely on 2D portrait images, bridging the gap between 2D talking faces and 3D avatars. Leveraging the rendering capability of 3DGS and the rich expressiveness of the expression latent space, the learned avatars can be reenacted to reproduce photo-realistic rendering images with subtle and accurate facial expressions. Experiments on an existing dataset and our newly proposed dataset offull-body talking avatars demonstrate the efficacy of our method. We also propose an audio-driven extension of our method with the help of 2D talking faces, opening new possibilities for interactive AI agents. Project page: https://initialneil.github.io/DEGAS.
Zhijing Shao, Duotun Wang, Qing-Yao Tian, Yao-Dong Yang, Hengyu Meng, Yu Zhang 0166, Kang Zhang 0001, Zeyu Wang 0003
3DV5
2025 HeadEvolver: Text to Head Avatars via Expressive and Attribute-Preserving Mesh Deformation
abstract
Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for generating stylized head avatars from text guidance, which leverages locally learnable mesh deformation and 2D diffusion priors to achieve high-quality digital assets for attribute-preserving manipulation. Given a template mesh, our method represents mesh deformation with perface Jacobians and adaptively modulates local deformation using a learnable vector field. This vector field enables anisotropic scaling while preserving the rotation of vertices, which can better express identity and geometric details. We also employ landmark- and contour-based regularization terms to balance the expressiveness and plausibility of generated head avatars from multiple views without relying on any specific shape prior. Our framework can generate realistic shapes and textures that can be further edited via text, while supporting seamless editing using the preserved attributes from the template mesh, such as 3DMM parameters, blendshapes, and UV coordinates. Extensive experiments demonstrate that our framework can generate diverse and expressive head avatars with high-quality meshes that artists can easily manipulate in 3D graphics software, facilitating downstream applications such as efficient asset creation and animation with preserved attributes.
Duotun Wang, Hengyu Meng, Zhijing Shao, Qianxi Liu, Lin Wang 0025, Mingming Fan 0001, Xiaohang Zhan, Zeyu Wang 0003
3DV2
2025 Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D Sculpting
abstract
Professional 3D asset creation often requires diverse sculpting brushes to add surface details and geometric structures. Despite recent progress in 3D generation, producing reusable sculpting brushes compatible with artists' workflows remains an open and challenging problem. These sculpting brushes are typically represented as vector displacement maps (VDMs), which existing models cannot easily generate compared to natural images. This paper presents Text2VDM, a novel framework for text-to-VDM brush generation through the deformation of a dense planar mesh guided by score distillation sampling (SDS). The original SDS loss is designed for generating full objects and struggles with generating desirable sub-object structures from scratch in brush generation. We refer to this issue as semantic coupling, which we address by introducing weighted blending of prompt tokens to SDS, resulting in a more accurate target distribution and semantic guidance. Experiments demonstrate that Text2VDM can generate diverse, high-quality VDM brushes for sculpting surface details and geometric structures. Our generated brushes can be seamlessly integrated into mainstream modeling software, enabling various applications such as mesh stylization and real-time interactive modeling.
Hengyu Meng, Duotun Wang, Zhijing Shao, Zeyu Wang 0003
ICCV1
2025 MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image Generation
abstract
Scroll images are a unique medium commonly used in virtual reality (VR) providing an immersive visual storytelling experience. Despite rapid advances in diffusion-based image generation, it remains an open research question to generate scroll images suitable for immersive, coherent, and controllable storytelling in VR. This paper proposes a multi-layered, diffusion-based scroll image generation framework with a novel semantic-aware denoising process. We incorporate layout prediction and style control modules to generate coherent scroll images of any aspect ratio. Based on the scroll image generation framework, we use different multi-window strategies to render diverse visual forms such as chains, rings, and forks for VR storytelling. Quantitative and qualitative evaluations demonstrate that our techniques can significantly enhance text-image consistency and visual coherence in scroll image generation, as well as the level of immersion and engagement of VR storytelling. We will release our source code to facilitate better collaborations on immersive storytelling between AI researchers and creative practitioners. https://magicscroll.github.io/
Bingyuan Wang, Hengyu Meng, Lanjiong Li, Yue Ma 0016, Qifeng Chen 0001, Zeyu Wang 0003
VR2
2024 Get Your Hands Dirty? A Comparative Study of Tool Usage and Perceptual Engagement in Physical and Digital Sculpting
abstract
The creation of 3D content, crucial in various applications, is often challenging and time-intensive. While digital tools are prevalent for 3D content creation, traditional clay sculpting offers an embodied experience that fosters artists’ perceptual engagement with physical space, enhancing their interactive and cognitive connection with the creation process. We conducted an eight-day live sculpting session at an art academy, systematically comparing the creative workflows of eight professional artists in both physical and digital mediums. Our qualitative and quantitative analysis include artists’ differences in tool usage between physical and digital sculpting, variations in visual and tactile perceptual engagement, and the potential for future integration of the two modalities. Our study provides insights into the benefits of physical and digital sculpting and may inform future design of hybrid interfaces for 3D content creation.
Hengyu Meng, Yanan Jin, Zeyu Wang 0003
Creativity & Cognition3