Xiuhua Jiang

dblp:39/3082 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-2422-942XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
YearPublicationVenuePosition
2026 GranuMamba: A multi-granularity state space model for co-speech gesture generation
Gaolin Yang, Jinhan Xie, Yifan Ge, Xiuhua Jiang, Jiangbo Xu
Expert Syst. Appl.5
2025 Toward Realistic Co-Speech Motion via Cross-Modal Spatial-Temporal Attention and Hand Memory Module
abstract
Co-speech video generation focuses on improving the authenticity of virtual characters by aligning their gestures and facial expressions with spoken audio. Despite recent advancements, existing methods often struggle with speech-gesture misalignment and unnatural hand motions. To address the issues, we propose a novel audio-driven gesture generation framework. This framework integrates a hierarchical diffusion model with multimodal feature disentanglement and dynamic fusion strategies. The core of our approach is the Cross-modal Spatial-Temporal Attention mechanism (CSTA), which ensures high-fidelity synchronization between audio and human motion while capturing fine-grained dynamics of hand and facial. By effectively disentangling different motion modalities, CSTA enhances the alignment between body part movements and the audio signal, leading to more natural and coherent video synthesis. Furthermore, to improve the physical plausibility and diversity of generated gestures, we introduce a Hand Memory Module (HMM). This module leverages a Vector Quantization-Variational Autoencoder (VQ-VAE) to learn a discrete gesture prior. By embedding these learned priors during the generation process, our method not only enhances temporal consistency but also preserves intricate details, mitigating common issues in prior work such as motion blur and detail loss. Experiments on the PATS and BEAT2 datasets demonstrate that CSTA surpasses existing methods in generating co-speech videos with more synchronized and natural hand motions, achieving state-of-the-art performance in both qualitative and quantitative evaluations. Project Page: CSTA-HMM
Dingwei Liu, Guibiao Liao, Xiuhua Jiang, Jiangbo Xu
ECAI4
2025 Co-speech video generation via motion transfer based on diffusion models
Zhiye Chen, Ruidi Zheng, Ruoyu Liu, Xiuhua Jiang
Neurocomputing5
2024 SR4KVQA: Video quality assessment database and metric for 4K super-resolution
Ruidi Zheng, Xiuhua Jiang
J. Vis. Commun. Image Represent.2
2024 Blind visual quality assessment for super-resolution images: database and model
Ruidi Zheng, Xiuhua Jiang, Hualong Yu
Multim. Tools Appl.2
2023 Learning a Practical SDR-to-HDRTV Up-conversion using New Dataset and Degradation Models
abstract
In media industry, the demand of SDR-to-HDRTV upconversion arises when users possess HDR-WCG (high dynamic range-wide color gamut) TVs while most off-the-shelf footage is still in SDR (standard dynamic range). The research community has started tackling this low-level vision task by learning-based approaches. When applied to real SDR, yet, current methods tend to produce dim and desaturated result, making nearly no improvement on viewing experience. Different from other network-oriented methods, we attribute such deficiency to training set (HDR-SDR pair). Consequently, we propose new HDRTV dataset (dubbed HDRTV4K) and new HDR-to-SDR degradation models. Then, it's used to train a luminance-segmented network (LSN) consisting of a global mapping trunk, and two Transformer branches on bright and dark luminance range. We also update assessment criteria by tailored metrics and subjective experiment. Finally, ablation studies are conducted to prove the effectiveness. Our work is available at: https://github.com/AndreGuo/HDRTVDM
Cheng Guo 0005, Leidong Fan, Xiuhua Jiang
CVPR4
2023 Blind image quality assessment for anchor-assisted adaptation to practical situations
Xiuhua Jiang
Multim. Tools Appl.2
2022 LHDR: HDR Reconstruction for Legacy Content Using a Lightweight DNN
Cheng Guo 0005, Xiuhua Jiang
ACCV (3)2
2020 Multimedia image quality assessment based on deep feature extraction
Xiuhua Jiang
Multim. Tools Appl.2
2007 Objective Perceptual Video Quality Measurement using a Foveation-Based Reduced Reference Algorithm
abstract
The properties and limitations of human visual system (HVS) are usually incorporated into objective video quality techniques for better correlation with subjective methods. In this paper, we propose a new objective assessment algorithm based on reduced reference video quality metrics (RR-VQM). The key point of our method is to design multilayer Spatial-Temporal (ST) regions based on human's fovea region, which can utilize the nonuniform resolution properties of HVS for high definition (HD) videos. Quality features are then extracted from ST regions for computing the video quality. Experimental results shows that this kind of multi-layer ST regions can improve the performance of objective methods, not only the results of evaluation, but can also accelerate the evaluation speed and reduce the amount of quality parameters to a large extent.
Fang Meng, Xiuhua Jiang
ICME2