Dengming Zhang

dblp:173/0915 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-6307-7692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Audio and music processing · 57% Visual content generation and editing · 30% Rendering · 8%
Artificial intelligence
2 papers
Generative modeling · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music generation
1.722025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Audio and music processing › music generation
video-to-music generation
1.722025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Machine learning › Generative modeling
flow matching
0.912025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Audio and music processing › music generation
controllable music generation
0.912025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Audio and music processing › music information retrieval
music emotion recognition
0.912025
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning · AAAI 2025
Visual content generation and editing
style control
0.912025
FonTS: Text Rendering with Typography and Style Controls · ICCV 2025
Rendering
text rendering
0.912025
FonTS: Text Rendering with Typography and Style Controls · ICCV 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
FonTS: Text Rendering with Typography and Style Controls · ICCV 2025
Visual content generation and editing
image generation
0.812024
StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation · UIST 2024
Visual content generation and editing
style transfer
0.812024
StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation · UIST 2024
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization
0.312025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Image and video processing › image sequence processing
temporal alignment
0.312025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212024
StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation · UIST 2024

Methods — techniques the papers use, named apart from their topics

user study · 1.5formative study · 1.5two-stage training · 0.9temporal alignment attention · 0.9style control adapter · 0.9spatial-temporal decomposition · 0.9parameter-efficient fine-tuning · 0.9meta-learning · 0.9flow-matching alignment · 0.9feature-free guidance · 0.9feature selection · 0.9feature alignment · 0.9dual-scale attention transformer · 0.9diffusion transformer · 0.9control-guided decoding · 0.9conditional fusion · 0.9
YearPublicationVenuePosition
2026 WordCon: Word-Level Typography Control in Visual Text Rendering
abstract
Visual text rendering represents a fundamental capability of large-scale text-to-image (T2I) models, yet achieving precise word-level controllability remains a significant challenge in this domain. While existing approaches primarily focus on text content accuracy, they often fail to provide fine-grained control over typographic attributes at the word level. To address this limitation, we introduce a comprehensive solution comprising three key components: (1) a novel word-level controlled scene text dataset and benchmark, (2) the Text-Image Alignment (TIA) framework that leverages cross-modal correspondence between textual queries and local image regions through grounding models, and (3) WordCon, a hybrid parameter-efficient fine-tuning (PEFT) method that employs selective parameter reparameterization to enhance both computational efficiency and model portability. The proposed framework incorporates additive supervision mechanisms: a masked loss at the latent level to focus on text regions, and a joint-attention loss at the feature level to promote disentanglement between different words. Extensive experimental evaluations demonstrate that our approach outperforms state-of-the-art methods in both qualitative and quantitative metrics. The proposed method exhibits remarkable versatility, enabling seamless integration with diverse pipelines, including artistic text rendering and image-conditioned text generation. Our datasets, source code, and models will be available for academic research.
Wenda Shi, Yiren Song, Zihan Rao, Dengming Zhang, Xingxing Zou
IEEE Trans. Circuits Syst. Video Technol.4
2025 Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
abstract
Dynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore, these methods often overlook the influence of individual differences on emotion perception, even though everyone has their own personalized emotional perception in the real world. Motivated by these issues, we explore more effective sequence processing methods and introduce the Personalized DMER (PDMER) problem, which requires models to predict emotions that align with personalized perception. Specifically, we propose a Dual-Scale Attention-Based Meta-Learning (DSAML) method. This method fuses features from a dual-scale feature extractor and captures both short and long-term dependencies using a dual-scale attention transformer, improving the performance in traditional DMER. To achieve PDMER, we design a novel task construction strategy that divides tasks by annotators. Samples in a task are annotated by the same annotator, ensuring consistent perception. Leveraging this strategy alongside meta-learning, DSAML can predict personalized perception of emotions with just one personalized annotation sample. Our objective and subjective experiments demonstrate that our method can achieve state-of-the-art performance in both traditional DMER and PDMER.
Dengming Zhang, Weitao You, Lingyun Sun, Pei Chen 0005
AAAI1
2025 FonTS: Text Rendering with Typography and Style Controls
abstract
Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show promise in automating these processes. However, these methods still encounter challenges like inconsistent fonts, style variation, and limited fine-grained control, particularly at the word-level. This paper proposes a two-stage DiT-based pipeline to address these problems by enhancing controllability over typography and style in text rendering. We introduce typography control fine-tuning (TC-FT), an parameter-efficient fine-tuning method (on $5\%$ key parameters) with enclosing typography control tokens (ETC-tokens), which enables precise word-level application of typographic features. To further address style inconsistency in text rendering, we propose a text-agnostic style control adapter (SCA) that prevents content leakage while enhancing style consistency. To implement TC-FT and SCA effectively, we incorporated HTML-render into the data synthesis pipeline and proposed the first word-level controllable dataset. Through comprehensive experiments, we demonstrate the effectiveness of our approach in achieving superior word-level typographic control, font consistency, and style consistency in text rendering tasks. The datasets and models will be available for academic use.
Wenda Shi, Yiren Song, Dengming Zhang, Xingxing Zou
ICCV3
2025 Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
abstract
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
Junxian Wu 0003, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen 0005, Lingyun Sun
ACM Multimedia4
2025 Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation
abstract
Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.
Weitao You, Heda Zuo, Junxian Wu 0003, Dengming Zhang, Zhibin Zhou 0002, Lingyun Sun
ACM Multimedia4
2024 StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation
abstract
Generative AI models have been widely used for image creation. However, generating images that are well-aligned with users’ personal styles on aesthetic features (e.g., color and texture) can be challenging due to the poor style expression and interpretation between humans and models. Through a formative study, we observed that participants showed a clear subjective perception of the desired style and variations in its strength, which directly inspired us to develop style-strength-based control and evaluation. Building on this, we present StyleFactory, an interactive system that helps users achieve style alignment. Our interface enables users to rank images based on their strengths in the desired style and visualizes the strength distribution of other images in that style from the model’s perspective. In this way, users can evaluate the understanding gap between themselves and the model, and define well-aligned personal styles for image creation through targeted iterations. Our technical evaluation and user study demonstrate that StyleFactory accurately generates images in specific styles, effectively facilitates style alignment in image creation workflow, stimulates creativity, and enhances the user experience in human-AI interactions.
Mingxu Zhou, Dengming Zhang, Weitao You, Chenghao Pan, Tianyu Lao, Pei Chen 0005
UIST2