Gaoge Han

dblp:365/5193 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0001-6731-8203ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.812024
HuTuMotion: Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback · AAAI 2024
Machine learning › Generative modeling
motion generation
0.812024
HuTuMotion: Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback · AAAI 2024
Human-AI interaction
human feedback
0.812024
HuTuMotion: Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback · AAAI 2024

Methods — techniques the papers use, named apart from their topics

latent diffusion model · 1.5human feedback optimization · 1.5
YearPublicationVenuePosition
2026 ReBaR: Reference-based reasoning for robust pose estimation from monocular images
Yongkang Cheng, Mingjiang Liang, Jifeng Ning, Gaoge Han, Wei Liu 0007, Shaoli Huang
Pattern Recognit.4
2026 Kernel-aware dynamic convolution for dense prediction
Gaoge Han, Mingjiang Liang, Jinglei Tang, Yongkang Cheng, Shaoli Huang, Wei Liu 0007
Pattern Recognit.1
2026 Context-assisted astrous deformable convolution for robust goat face detection and identification
Gaoge Han, Lianyue Zhang, Zihan Bai, Ruizi Han, Jinglei Tang
Vis. Comput.1
2025 Inter - Diffusion Generation Model of Speakers and Listeners for Effective Communication
abstract
Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.
Jinhe Huang, Yongkang Cheng, Minghang Yu, Gaoge Han, Jinwei Li 0003, Jing Zhang 0164, Shilei Wang 0001, Xingjian Gu
ICMR4
2025 Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
abstract
Audio-driven simultaneous gesture generation is vital for human-computer communication, AI games, and film pro-duction. While previous research has shown promise, there are still limitations. Methods based on VAEs are accompa-nied by issues of local jitter and global instability, whereas methods based on diffusion models are hampered by low generation efficiency. This is because the denoising process of DDPM in the latter relies on the assumption that the noise added at each step is sampled from a unimodal distribution, and the noise values are small. DDIM bor-rows the idea from the Euler method for solving differential equations, disrupts the Markov chain process, and increases the noise step size to reduce the number of denoising steps, thereby accelerating generation. However, simply increasing the step size during the step-by-step denoising process causes the results to gradually deviate from the original data distribution, leading to a significant drop in the quality of the generated actions and the emergence of unnatural artifacts. In this paper, we break the assumptions of DDPM and achieves breakthrough progress in denoising speed and fidelity. Specifically, we introduce a conditional GAN to capture audio control signals and implicitly match the multimodal denoising distribution between the diffusion and denoising steps within the same sampling step, aiming to sample larger noise values and apply fewer denoising steps for high-speed generation. In addition, to enable the model to generate high-fidelity global gestures and avoid artifacts, we introduce an explicit motion geometric loss to enhance the quality and global stability of the generated gestures. Numerous qualitative and quantitative experiments show that compared to contemporary diffusion-based methods, our method offers faster generation speed and higher fidelity, and compared to non-diffusion methods, it provides a more stable global effect and a more natural user experience.
Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Gaoge Han, Jifeng Ning, Wei Liu 0007
WACV4
2025 ReinDiffuse: Crafting Physically Plausible Motions with Reinforced Diffusion Model
abstract
Generating human motion from textual descriptions is a challenging task. Existing methods either struggle with physical credibility or are limited by the complexities of physics simulations. In this paper, we present ReinDiffuse that combines reinforcement learning with motion diffusion model to generate physically credible human motions that align with textual descriptions. Our method adapts Motion Diffusion Model to output a parameterized distribution of actions, making them compatible with reinforcement learning paradigms. We employ reinforcement learning with the objective of maximizing physically plausible rewards to optimize motion generation for physical fidelity. Our approach outperforms existing state-of-the-art models on two major datasets, HumanML3D and KIT-ML, achieving significant improvements in physical plausibility and motion quality. Project: https://reindiffuse.github.io/
Gaoge Han, Mingjiang Liang, Jinglei Tang, Yongkang Cheng, Wei Liu 0007, Shaoli Huang
WACV1
2024 HuTuMotion: Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback
abstract
We introduce HuTuMotion, an innovative approach for generating natural human motions that navigates latent motion diffusion models by leveraging few-shot human feedback. Unlike existing approaches that sample latent variables from a standard normal prior distribution, our method adapts the prior distribution to better suit the characteristics of the data, as indicated by human feedback, thus enhancing the quality of motion generation. Furthermore, our findings reveal that utilizing few-shot feedback can yield performance levels on par with those attained through extensive human feedback. This discovery emphasizes the potential and efficiency of incorporating few-shot human-guided optimization within latent diffusion models for personalized and style-aware human motion generation applications. The experimental results show the significantly superior performance of our method over existing state-of-the-art approaches.
Gaoge Han, Shaoli Huang, Mingming Gong, Jinglei Tang
AAAI1
2024 SIAM: A parameter-free, Spatial Intersection Attention Module
Gaoge Han, Shaoli Huang, Fang Zhao 0006, Jinglei Tang
Pattern Recognit.1