EDBT 2026 Demo / reviewers in the wild / expert
Huangbiao Xu
dblp:376/3477
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-3717-8713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentabstractMultimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities are frequently unavailable at the inference stage in reality. The absence of any modality often renders existing multimodal models inoperable. Furthermore, it triggers catastrophic performance degradation due to interruptions in cross-modal interactions. To address this issue, we propose a novel Missing Completion Framework with Mixture of Experts (MCMoE) that unifies unimodal and joint representation learning in single-stage training. Specifically, we propose an adaptive gated modality generator that dynamically fuses available information to reconstruct missing modalities. We then design modality experts to learn unimodal knowledge and dynamically mix the knowledge of all experts to extract cross-modal joint representations. With a mixture of experts, missing modalities are further refined and complemented. Finally, in the training phase, we mine the complete multimodal features and unimodal expert knowledge to guide modality generation and generation-based joint representation extraction. Extensive experiments demonstrate that our MCMoE achieves state-of-the-art results in both complete and incomplete multimodal learning on three public AQA benchmarks. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Rui Xu 0028, Jinglin Xu |
AAAI | 1 |
| 2026 | Integrating perceptual cues with mixture-of-experts for low-light image restoration
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Hui Da, Wenxi Liu, Lifang Wei |
Neural Networks | 3 |
| 2025 | DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseabstractThe fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix proposes a bidirectional spatial-temporal context optical flow correction (BOFC) method. This approach leverages the consistency and complementarity of motion information between two modalities: optical flow, which excels at pixel capture, and lightweight skeleton data. It enables the extraction of pixel-level motion changes and the correction of abnormal skeleton data. Furthermore, we propose a part-level dance dataset (Dancer Parts) and part-level motion feature extraction based on task decoupling (PETD). This aims to decouple complex whole-body parts tracking into fine-grained limb-level motion extraction, enhancing the confidence of temporal information and the accuracy of correction for abnormal data. Finally, we present the DNV dataset, which simulates fully neat group dance scenes and provides reliable labels and validation methods for the newly introduced group dance neatness assessment (GDNA). To the best of our knowledge, this is the first work to develop quantitative criteria for assessing limb and joint neatness in group dance. We conduct experiments on DNV and video-based public JHMDB datasets. Our method effectively corrects abnormal skeleton points, flexibly embeds, and improves the accuracy of existing pose estimation algorithms. Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Peirong Xu, Wenzhong Guo |
AAAI | 1 |
| 2025 | URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image RestorationabstractExisting low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexible and effective degradation restoration for low-light images. Specifically, we customize the core URWKV block to perceive and analyze complex degradations by leveraging multiple intra- and inter-stage states. First, inspired by the pupil mechanism in the human visual system, we propose Luminance-adaptive Normalization (LAN) that adjusts normalization parameters based on rich inter-stage states, allowing for adaptive, scene-aware luminance modulation. Second, we aggregate multiple intra-stage states through exponential moving average approach, effectively capturing subtle variations while mitigating information loss inherent in the single-state mechanism. To reduce the degradation effects commonly associated with conventional skip connections, we propose the State-aware Selective Fusion (SSF) module, which dynamically aligns and integrates multi-state features across encoder stages, selectively fusing contextual information. In comparison to state-of-the-art models, our URWKV model achieves superior performance on various benchmarks, while requiring significantly fewer parameters and computational resources. Code is available at: https://github.com/FZU-N/URWKV. Rui Xu 0028, Yuzhen Niu, Yuezhou Li, Huangbiao Xu, Wenxi Liu, Yuzhong Chen 0001 |
CVPR | 4 |
| 2025 | Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentabstractLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large number of model parameters to learn potential associations between actions and music. To address this issue, we propose a language-guided audio-visual learning (MLAVL) framework that models "audio-action-visual" correlations guided by low-cost language modality. In our framework, multidimensional domain-based actions form action knowledge graphs, motivating audio-visual modalities to focus on task-relevant actions. We further design a shared-specific context encoder to integrate deep multimodal semantics, and an audio-visual cross-modal fusion module to evaluate action-music consistency. To match the sport’s rules, we then propose a dual-branch prompt-guided grading module to weigh both visual and audio-visual performance. Extensive experiments demonstrate that our approach achieves state-of-the-art on four public long-term sports benchmarks while maintaining low parameters.1 Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Wenzhong Guo |
CVPR | 1 |
| 2025 | Progressive Modality-Adaptive Interactive Network for Multi-Modality Image FusionabstractMulti-modality image fusion (MMIF) integrates features from distinct modalities to enhance visual quality and improve downstream task performance. However, existing methods often overlook the sparsity variations and dynamic correlations between infrared and visible images, potentially limiting the utilization of both modalities. To address these challenges, we propose the Progressive Modality-Adaptive Interactive Network (PoMAI), a novel framework that not only dynamically adapts to the sparsity and structural disparities of each modality but also enhances inter-modal correlations, thereby optimizing fusion quality. The training process consists of two stages: in the first stage, the Neighbor-Group Matching Model (NGMM) models the high sparsity of infrared features, while the Context-Aware Modeling Network (CAMN) captures rich structural details in visible features, jointly refining modality-specific characteristics for fusion. In the second stage, the Modality-Interactive Compensation Module (MICM) refines inter-modal correlations via dynamic compensation mechanism, while freezing the first-stage modules to focus MICM solely on the compensation task. Extensive experiments on benchmark datasets demonstrate that PoMAI surpasses state-of-the-art methods in fusion quality and excels in downstream tasks. Chaowei Huang, Yaru Su, Huangbiao Xu, Xiao Ke |
IJCAI | 3 |
| 2025 | The Devil in the Stego Image: Far from Being Usable in Real-World ScenariosabstractDigital images, serving as the primary carrier of information, have been wildly spread on the Internet. Image steganography is a technology that employs images as the carrier for information hiding. While current deep image steganography demonstrated impressive encoding abilities across various media, two serious problems have been overlooked in deep image-to-image steganography and hinder its application under real-world scenarios, which we define as the problem of Pixel Value Overflow and Gap of Precision. In this paper, we explore the cause of those problems and introduce a plug-and-play Universal Suppressor to solve the application problems of deep image-to-image steganography in real-world scenarios, which can be flexibly applied to various models with different structures. Experiments demonstrate that our Universal Suppressor performs well in existing state-of-the-art (SOTA) models and confers them with intrinsic robustness for real-world deployment. The code will be released at https://github.com/aoli-gei/USP. Huanqi Wu 0001, Huangbiao Xu, Xiao Ke |
ACM Multimedia | 2 |
| 2025 | IPCMoE: Integrating Perceptual Cues with Mixture-of-Experts for Joint Low-Light Image Enhancement and DeblurringabstractVisual perception of nighttime images is often compromised by co-existing low-light and blur degradations. While recent methods have made progress in jointly solving these degradations, the diversity of patterns and intensities in degradation has not been properly considered, leading to inconsistent illumination and unintended artifacts. In response, we propose to integrate perceptual cues with mixture-of-experts (IPCMoE) to achieve flexible processing for low-light blurry images. By exploiting the perceptual cues, we strategically combine dedicated experts with the selective collaboration approach for feature enlightening and texture restoration. To this end, we develop perceptual-integrated MoEs by designing customized routers and task-depended experts. Specifically, the texture memorial MoE is developed to preserve valuable features to restore high-fidelity details, and the enhancement MoE that adaptively integrates enlightening cues and texture cues is designed to formulate the relationship between feature enlightening and texture restoration, thereby achieving dynamic image processing. Extensive experiments show that our method achieves state-of-the-art performance on LOL-Blur and Real-LOL-Blur datasets. Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Hui Da, Rui Xu 0028, Wenxi Liu |
ACM Multimedia | 3 |
| 2025 | Collaboratively enhanced and integrated detail-context information for low-light image enhancement
Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Yuzhong Chen 0001 |
Pattern Recognit. | 3 |
| 2025 | MSP: Multimodal Self-Attention Prompt LearningabstractMultimodal prompt learning has emerged as an effective strategy for adapting vision-language models such as CLIP to downstream tasks. However, conventional approaches typically operate at the input level, forcing learned prompts to propagate through a sequence of frozen Transformer layers. This indirect adaptation introduces cumulative geometric distortions, a limitation that we formalize as the indirect learning dilemma (ILD), leading to overfitting of the base class and reduced generalization to novel classes. To overcome this challenge, we propose the Multimodal Self-Attention Prompt (MSP) framework, which shifts adaptation into the semantic core of the model by injecting learnable prompts directly into the key and value sequences of attention blocks. This direct modulation preserves the pretrained embedding geometry while enabling more precise downstream adaptation. MSP further incorporates distance-aware optimization to maintain semantic consistency with CLIP's original representation space, and partial prompt learning via stochastic dimension masking to improve robustness and prevent over-specialization. Extensive evaluations across 11 benchmarks demonstrate the effectiveness of MSP. It achieves a state-of-the-art harmonic mean accuracy of 80.67%, with 77.32% accuracy on novel classes-representing a 2.18% absolute improvement over prior methods-while requiring only 0.11M learnable parameters. Notably, MSP surpasses CLIP's zero-shot performance on 10 out of 11 datasets, establishing a new paradigm for efficient and generalizable prompt-based adaptation. Our implementation is available at https://github.com/laixinyi023/Multimodal-Self-Attention-Prompt. Xinyi Lai, Xiao Ke, Huangbiao Xu, Shanghui Wu, Wenzhong Guo |
IEEE Trans. Image Process. | 3 |
| 2025 | Quality-Guided Vision-Language Learning for Long-Term Action Quality AssessmentabstractLong-term action quality assessment poses a challenging visual task since it requires assessing technical actions at different skill levels in a long video. Recent state-of-the-art methods incorporate additional modality information to aid in understanding action semantics, which incurs extra annotation costs and imposes higher constraints on action scenes and datasets. To address this issue, we propose a Quality-Guided Vision-Language Learning (QGVL) method to map visual features into appropriate fine-grained intervals of quality scores. Specifically, we use a set of quality-related textual prompts as quality prototypes to guide the discrimination and aggregation of specific visual actions. To avoid fuzzy rule mapping, we further propose a progressive semantic learning strategy with a Granularity-Adaptive Semantic Learning Module (GSLM) that refines accurate score intervals from coarse to fine at clip, grade, and score levels. The quality-related semantics we designed are universal to all types of action scenarios without any additional annotations. Extensive experiments show that our approach outperforms previous work by a significant margin and establishes new state-of-the-art on four public AQA benchmarks: Rhythmic Gymnastics, Fis-V, FS1000, and FineFS. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Yuezhou Li, Rui Xu 0028, Wenzhong Guo |
IEEE Trans. Multim. | 1 |
| 2024 | Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment
Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu 0028, Huanqi Wu 0001, Wenzhong Guo |
ECCV (42) | 1 |
| 2024 | Two-path target-aware contrastive regression for action quality assessment
Xiao Ke, Huangbiao Xu, Wenzhong Guo |
Inf. Sci. | 2 |
| 2024 | Bilateral Interaction for Local-Global Collaborative Perception in Low-Light Image EnhancementabstractLow-light image enhancement is a challenging task due to the limited visibility in dark environments. While recent advances have shown progress in integrating CNNs and Transformers, the inadequate local-global perceptual interactions still impedes their application in complex degradation scenarios. To tackle this issue, we propose BiFormer, a lightweight framework that facilitates local-global collaborative perception via bilateral interaction. Specifically, our framework introduces a core CNN-Transformer collaborative perception block (CPB) that combines local-aware convolutional attention (LCA) and global-aware recursive Transformer (GRT) to simultaneously preserve local details and ensure global consistency. To promote perceptual interaction, we adopt bilateral interaction strategy for both local and global perception, which involves local-to-global second-order interaction (SoI) in the dual-domain, as well as a mixed-channel fusion (MCF) module for global-to-local interaction. The MCF is also a highly efficient feature fusion module tailored for degraded features. Extensive experiments conducted on low-level and high-level tasks demonstrate that BiFormer achieves state-of-the-art performance. Furthermore, it exhibits a significant reduction in model parameters and computational cost compared to existing Transformer-based low-light image enhancement methods. Rui Xu 0028, Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Yuzhong Chen 0001, Tiesong Zhao |
IEEE Trans. Multim. | 4 |