VLDB 2026 Research / reviewers in the wild / expert
Xiuhua Jiang
dblp:39/3082
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-2422-942XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GranuMamba: A multi-granularity state space model for co-speech gesture generation
Gaolin Yang, Jinhan Xie, Yifan Ge, Xiuhua Jiang, Jiangbo Xu |
Expert Syst. Appl. | 5 |
| 2025 | Toward Realistic Co-Speech Motion via Cross-Modal Spatial-Temporal Attention and Hand Memory ModuleabstractCo-speech video generation focuses on improving the authenticity of virtual characters by aligning their gestures and facial expressions with spoken audio. Despite recent advancements, existing methods often struggle with speech-gesture misalignment and unnatural hand motions. To address the issues, we propose a novel audio-driven gesture generation framework. This framework integrates a hierarchical diffusion model with multimodal feature disentanglement and dynamic fusion strategies. The core of our approach is the Cross-modal Spatial-Temporal Attention mechanism (CSTA), which ensures high-fidelity synchronization between audio and human motion while capturing fine-grained dynamics of hand and facial. By effectively disentangling different motion modalities, CSTA enhances the alignment between body part movements and the audio signal, leading to more natural and coherent video synthesis. Furthermore, to improve the physical plausibility and diversity of generated gestures, we introduce a Hand Memory Module (HMM). This module leverages a Vector Quantization-Variational Autoencoder (VQ-VAE) to learn a discrete gesture prior. By embedding these learned priors during the generation process, our method not only enhances temporal consistency but also preserves intricate details, mitigating common issues in prior work such as motion blur and detail loss. Experiments on the PATS and BEAT2 datasets demonstrate that CSTA surpasses existing methods in generating co-speech videos with more synchronized and natural hand motions, achieving state-of-the-art performance in both qualitative and quantitative evaluations. Project Page: CSTA-HMM Dingwei Liu, Guibiao Liao, Xiuhua Jiang, Jiangbo Xu |
ECAI | 4 |
| 2025 | Co-speech video generation via motion transfer based on diffusion models
Zhiye Chen, Ruidi Zheng, Ruoyu Liu, Xiuhua Jiang |
Neurocomputing | 5 |
| 2024 | SR4KVQA: Video quality assessment database and metric for 4K super-resolution
Ruidi Zheng, Xiuhua Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Blind visual quality assessment for super-resolution images: database and model
Ruidi Zheng, Xiuhua Jiang, Hualong Yu |
Multim. Tools Appl. | 2 |
| 2023 | Learning a Practical SDR-to-HDRTV Up-conversion using New Dataset and Degradation ModelsabstractIn media industry, the demand of SDR-to-HDRTV upconversion arises when users possess HDR-WCG (high dynamic range-wide color gamut) TVs while most off-the-shelf footage is still in SDR (standard dynamic range). The research community has started tackling this low-level vision task by learning-based approaches. When applied to real SDR, yet, current methods tend to produce dim and desaturated result, making nearly no improvement on viewing experience. Different from other network-oriented methods, we attribute such deficiency to training set (HDR-SDR pair). Consequently, we propose new HDRTV dataset (dubbed HDRTV4K) and new HDR-to-SDR degradation models. Then, it's used to train a luminance-segmented network (LSN) consisting of a global mapping trunk, and two Transformer branches on bright and dark luminance range. We also update assessment criteria by tailored metrics and subjective experiment. Finally, ablation studies are conducted to prove the effectiveness. Our work is available at: https://github.com/AndreGuo/HDRTVDM Cheng Guo 0005, Leidong Fan, Xiuhua Jiang |
CVPR | 4 |
| 2023 | Blind image quality assessment for anchor-assisted adaptation to practical situations
Xiuhua Jiang |
Multim. Tools Appl. | 2 |
| 2022 | LHDR: HDR Reconstruction for Legacy Content Using a Lightweight DNN
Cheng Guo 0005, Xiuhua Jiang |
ACCV (3) | 2 |
| 2020 | Multimedia image quality assessment based on deep feature extraction
Xiuhua Jiang |
Multim. Tools Appl. | 2 |
| 2007 | Objective Perceptual Video Quality Measurement using a Foveation-Based Reduced Reference AlgorithmabstractThe properties and limitations of human visual system (HVS) are usually incorporated into objective video quality techniques for better correlation with subjective methods. In this paper, we propose a new objective assessment algorithm based on reduced reference video quality metrics (RR-VQM). The key point of our method is to design multilayer Spatial-Temporal (ST) regions based on human's fovea region, which can utilize the nonuniform resolution properties of HVS for high definition (HD) videos. Quality features are then extracted from ST regions for computing the video quality. Experimental results shows that this kind of multi-layer ST regions can improve the performance of objective methods, not only the results of evaluation, but can also accelerate the evaluation speed and reduce the amount of quality parameters to a large extent. Fang Meng, Xiuhua Jiang |
ICME | 2 |