VLDB 2026 Research / reviewers in the wild / expert
Huanqi Wu 0001
dblp:288/6796-1
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0008-4518-3273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentabstractMultimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities are frequently unavailable at the inference stage in reality. The absence of any modality often renders existing multimodal models inoperable. Furthermore, it triggers catastrophic performance degradation due to interruptions in cross-modal interactions. To address this issue, we propose a novel Missing Completion Framework with Mixture of Experts (MCMoE) that unifies unimodal and joint representation learning in single-stage training. Specifically, we propose an adaptive gated modality generator that dynamically fuses available information to reconstruct missing modalities. We then design modality experts to learn unimodal knowledge and dynamically mix the knowledge of all experts to extract cross-modal joint representations. With a mixture of experts, missing modalities are further refined and complemented. Finally, in the training phase, we mine the complete multimodal features and unimodal expert knowledge to guide modality generation and generation-based joint representation extraction. Extensive experiments demonstrate that our MCMoE achieves state-of-the-art results in both complete and incomplete multimodal learning on three public AQA benchmarks. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Rui Xu 0028, Jinglin Xu |
AAAI | 2 |
| 2026 | MDANet: A Lightweight Multi-Task Dynamic Adaptive Network for Real-Time Visual Perception in Autonomous Driving
Xiao Ke, Jingyi Fang, Chaoying Chen, Huanqi Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseabstractThe fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix proposes a bidirectional spatial-temporal context optical flow correction (BOFC) method. This approach leverages the consistency and complementarity of motion information between two modalities: optical flow, which excels at pixel capture, and lightweight skeleton data. It enables the extraction of pixel-level motion changes and the correction of abnormal skeleton data. Furthermore, we propose a part-level dance dataset (Dancer Parts) and part-level motion feature extraction based on task decoupling (PETD). This aims to decouple complex whole-body parts tracking into fine-grained limb-level motion extraction, enhancing the confidence of temporal information and the accuracy of correction for abnormal data. Finally, we present the DNV dataset, which simulates fully neat group dance scenes and provides reliable labels and validation methods for the newly introduced group dance neatness assessment (GDNA). To the best of our knowledge, this is the first work to develop quantitative criteria for assessing limb and joint neatness in group dance. We conduct experiments on DNV and video-based public JHMDB datasets. Our method effectively corrects abnormal skeleton points, flexibly embeds, and improves the accuracy of existing pose estimation algorithms. Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Peirong Xu, Wenzhong Guo |
AAAI | 3 |
| 2025 | Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentabstractLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large number of model parameters to learn potential associations between actions and music. To address this issue, we propose a language-guided audio-visual learning (MLAVL) framework that models "audio-action-visual" correlations guided by low-cost language modality. In our framework, multidimensional domain-based actions form action knowledge graphs, motivating audio-visual modalities to focus on task-relevant actions. We further design a shared-specific context encoder to integrate deep multimodal semantics, and an audio-visual cross-modal fusion module to evaluate action-music consistency. To match the sport’s rules, we then propose a dual-branch prompt-guided grading module to weigh both visual and audio-visual performance. Extensive experiments demonstrate that our approach achieves state-of-the-art on four public long-term sports benchmarks while maintaining low parameters.1 Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Wenzhong Guo |
CVPR | 3 |
| 2025 | The Devil in the Stego Image: Far from Being Usable in Real-World ScenariosabstractDigital images, serving as the primary carrier of information, have been wildly spread on the Internet. Image steganography is a technology that employs images as the carrier for information hiding. While current deep image steganography demonstrated impressive encoding abilities across various media, two serious problems have been overlooked in deep image-to-image steganography and hinder its application under real-world scenarios, which we define as the problem of Pixel Value Overflow and Gap of Precision. In this paper, we explore the cause of those problems and introduce a plug-and-play Universal Suppressor to solve the application problems of deep image-to-image steganography in real-world scenarios, which can be flexibly applied to various models with different structures. Experiments demonstrate that our Universal Suppressor performs well in existing state-of-the-art (SOTA) models and confers them with intrinsic robustness for real-world deployment. The code will be released at https://github.com/aoli-gei/USP. Huanqi Wu 0001, Huangbiao Xu, Xiao Ke |
ACM Multimedia | 1 |
| 2025 | Quality-Guided Vision-Language Learning for Long-Term Action Quality AssessmentabstractLong-term action quality assessment poses a challenging visual task since it requires assessing technical actions at different skill levels in a long video. Recent state-of-the-art methods incorporate additional modality information to aid in understanding action semantics, which incurs extra annotation costs and imposes higher constraints on action scenes and datasets. To address this issue, we propose a Quality-Guided Vision-Language Learning (QGVL) method to map visual features into appropriate fine-grained intervals of quality scores. Specifically, we use a set of quality-related textual prompts as quality prototypes to guide the discrimination and aggregation of specific visual actions. To avoid fuzzy rule mapping, we further propose a progressive semantic learning strategy with a Granularity-Adaptive Semantic Learning Module (GSLM) that refines accurate score intervals from coarse to fine at clip, grade, and score levels. The quality-related semantics we designed are universal to all types of action scenarios without any additional annotations. Extensive experiments show that our approach outperforms previous work by a significant margin and establishes new state-of-the-art on four public AQA benchmarks: Rhythmic Gymnastics, Fis-V, FS1000, and FineFS. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Yuezhou Li, Rui Xu 0028, Wenzhong Guo |
IEEE Trans. Multim. | 2 |
| 2024 | StegFormer: Rebuilding the Glory of Autoencoder-Based SteganographyabstractImage hiding aims to conceal one or more secret images within a cover image of the same resolution. Due to strict capacity requirements, image hiding is commonly called large-capacity steganography. In this paper, we propose StegFormer, a novel autoencoder-based image-hiding model. StegFormer can conceal one or multiple secret images within a cover image of the same resolution while preserving the high visual quality of the stego image. In addition, to mitigate the limitations of current steganographic models in real-world scenarios, we propose a normalizing training strategy and a restrict loss to improve the reliability of the steganographic models under realistic conditions. Furthermore, we propose an efficient steganographic capacity expansion method to increase the capacity of steganography and enhance the efficiency of secret communication. Through this approach, we can increase the relative payload of StegFormer to 96 bits per pixel without any training strategy modifications. Experiments demonstrate that our StegFormer outperforms existing state-of-the-art (SOTA) models. In the case of single-image steganography, there is an improvement of more than 3 dB and 5 dB in PSNR for secret/recovery image pairs and cover/stego image pairs. Xiao Ke, Huanqi Wu 0001, Wenzhong Guo |
AAAI | 2 |
| 2024 | Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment
Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu 0028, Huanqi Wu 0001, Wenzhong Guo |
ECCV (42) | 5 |