Muhammad Hamza Mughal

dblp:287/2063 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 46% Language models and text generation · 41% Speech recognition and synthesis · 14%
Computer graphics and multimedia
4 papers
Computer animation and physical simulation · 91% Audio and music processing · 9%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.332025
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis · CVPR 2025
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis · CVPR 2024
MoFusion: A Framework for Denoising-Diffusion-Based Motion Synthesis · CVPR 2023
Computer animation and physical simulation › gesture generation
co-speech gesture generation
1.622025
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis · CVPR 2025
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis · CVPR 2024
Computer animation and physical simulation
gesture generation
1.622025
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis · CVPR 2025
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis · CVPR 2024
Natural language and speech › Language models and text generation › natural language understanding
discourse modeling
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis · CVPR 2025
Natural language and speech › Speech recognition and synthesis
speech language model
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Machine learning › Generative modeling › motion generation
conditional motion generation
0.712023
MoFusion: A Framework for Denoising-Diffusion-Based Motion Synthesis · CVPR 2023
Computer animation and physical simulation › motion synthesis › human motion synthesis
diffusion-based motion generation
0.712023
MoFusion: A Framework for Denoising-Diffusion-Based Motion Synthesis · CVPR 2023
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.712023
MoFusion: A Framework for Denoising-Diffusion-Based Motion Synthesis · CVPR 2023
Information retrieval
retrieval-augmented generation
0.312025
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis · CVPR 2025
Audio and music processing
speech processing
0.312025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Computer animation and physical simulation › audio-driven animation
speech-driven animation
0.212024
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis · CVPR 2024

Methods — techniques the papers use, named apart from their topics

diffusion model · 4.1retrieval guidance · 2.6DDIM inversion · 2.6text infilling · 1.7feature alignment · 1.7VQ-VAE · 1.7guidance objectives · 1.5kinematic losses · 1.3denoising diffusion · 1.3
YearPublicationVenuePosition
2025 Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
abstract
Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse.For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse.In this work, we investigate whether the joint modeling of gestures using human motion sequences and language can improve spoken discourse modeling in language models.To integrate gestures into language models, we first encode 3D human motion sequences into discrete gesture tokens using a VQ-VAE.These gesture token embeddings are then aligned with text embeddings through feature alignment, mapping them into the text embedding space.To evaluate the gesture-aligned language model on spoken discourse, we construct text infilling tasks targeting three key discourse cues grounded in linguistic research: discourse connectives, stance markers, and quantifiers.Results show that incorporating gestures enhances marker prediction accuracy across the three tasks, highlighting the complementary information that gestures can offer in modeling spoken discourse.We view this work as an initial step toward leveraging non-verbal cues to advance spoken language modeling in language models.
Varsha Suresh, Muhammad Hamza Mughal, Christian Theobalt, Vera Demberg
ACL (1)2
2025 Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
abstract
Non-Verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-GESTURE, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.
Muhammad Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg, Christian Theobalt
CVPR1
2024 ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
abstract
Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal, semantically coherent gestures require modeling the complex interactions between the language and human motion, and can be controlled by focusing on certain words. Therefore, we present ConvoFusion, a diffusion-based approach for multi-modal gesture synthesis, which can not only generate gestures based on multi-modal speech inputs, but can also facilitate controllability in gesture synthesis. Our method proposes two guidance objectives that allow the users to modulate the impact of different conditioning modalities (e.g. audio vs text) as well as to choose certain words to be emphasized during gesturing. Our method is versatile in that it can be trained either for generating monologue gestures or even the conversational gestures. To further advance the research on multi-party interactive gestures, the DndGroup Gesture dataset is released, which contains 6 hours of gesture data showing 5 people interacting with one another. We compare our method with several recent works and demonstrate effectiveness of our method on a variety of tasks. We urge the reader to watch our supplementary video at our webpage.
Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli, Marc Habermann, Christian Theobalt
CVPR1
2023 MoFusion: A Framework for Denoising-Diffusion-Based Motion Synthesis
abstract
Conventional methods for human motion synthesis have either been deterministic or have had to struggle with the trade-off between motion diversity vs motion quality. In response to these limitations, we introduce MoFusion, i.e., a new denoising-diffusion-based framework for high-quality conditional human motion synthesis that can synthesise long, temporally plausible, and semantically accurate motions based on a range of conditioning contexts (such as music and text). We also present ways to introduce well-known kinematic losses for motion plausibility within the motion-diffusion framework through our scheduled weighting strategy. The learned latent space can be used for several interactive motion-editing applications like in-betweening, seed-conditioning, and text-based editing, thus, providing crucial abilities for virtual-character animation and robotics. Through comprehensive quantitative evaluations and a perceptual user study, we demonstrate the effectiveness of MoFusion compared to the state of the art on established benchmarks in the literature. We urge the reader to watch our supplementary video at https://vcai.mpi-inf.mpg.de/projects/MoFusion/.
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, Christian Theobalt
CVPR2