VLDB 2026 Research / reviewers in the wild / expert
Rui Xu 0028
dblp:00/4859-28
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-3767-6816ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentabstractMultimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities are frequently unavailable at the inference stage in reality. The absence of any modality often renders existing multimodal models inoperable. Furthermore, it triggers catastrophic performance degradation due to interruptions in cross-modal interactions. To address this issue, we propose a novel Missing Completion Framework with Mixture of Experts (MCMoE) that unifies unimodal and joint representation learning in single-stage training. Specifically, we propose an adaptive gated modality generator that dynamically fuses available information to reconstruct missing modalities. We then design modality experts to learn unimodal knowledge and dynamically mix the knowledge of all experts to extract cross-modal joint representations. With a mixture of experts, missing modalities are further refined and complemented. Finally, in the training phase, we mine the complete multimodal features and unimodal expert knowledge to guide modality generation and generation-based joint representation extraction. Extensive experiments demonstrate that our MCMoE achieves state-of-the-art results in both complete and incomplete multimodal learning on three public AQA benchmarks. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Rui Xu 0028, Jinglin Xu |
AAAI | 5 |
| 2026 | Integrating perceptual cues with mixture-of-experts for low-light image restoration
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Hui Da, Wenxi Liu, Lifang Wei |
Neural Networks | 4 |
| 2026 | ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T TrackingabstractRGB-T tracking benefits from the complementary nature of RGB and TIR modalities, yet their relative reliability for target localization often shifts over time. Most existing trackers fail to adapt to such modality and temporal dynamics in a unified and effective manner, resulting in target representations that are neither discriminative nor temporally consistent. In this paper, we propose ProMoT, a novel tracking framework that jointly integrates cross-modal and temporal cues into a progressive prompting process, enabling continuous retrieval of target-aware representations. Specifically, we design an adaptive target query generator (QueryGen), which selectively aggregates informative spatio-temporal cues from diverse ghost representations through the dynamic sparse ghost fusion mechanism, thereby enabling the generation of target-aware queries. To further preserve fine-grained, temporally consistent target cues, we introduce a high-order contextual prompt updater (PromptUpdater), which encodes high-order cross-modal representations from current and previous frames. These prompts establish the compact and discriminative inter-frame context to not only refine the current frame’s features but also guide target localization in future frames. All components are built upon a parameter-shared backbone for RGB and TIR inputs, forming our complete ProMoT framework. Extensive experiments on both complete and missing modality RGB-T tracking benchmarks show that ProMoT consistently achieves state-of-the-art performance while balancing efficiency. Rui Xu 0028, Si Chen 0002, Yuzhen Niu, Yan Yan 0001, Dahan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseabstractThe fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix proposes a bidirectional spatial-temporal context optical flow correction (BOFC) method. This approach leverages the consistency and complementarity of motion information between two modalities: optical flow, which excels at pixel capture, and lightweight skeleton data. It enables the extraction of pixel-level motion changes and the correction of abnormal skeleton data. Furthermore, we propose a part-level dance dataset (Dancer Parts) and part-level motion feature extraction based on task decoupling (PETD). This aims to decouple complex whole-body parts tracking into fine-grained limb-level motion extraction, enhancing the confidence of temporal information and the accuracy of correction for abnormal data. Finally, we present the DNV dataset, which simulates fully neat group dance scenes and provides reliable labels and validation methods for the newly introduced group dance neatness assessment (GDNA). To the best of our knowledge, this is the first work to develop quantitative criteria for assessing limb and joint neatness in group dance. We conduct experiments on DNV and video-based public JHMDB datasets. Our method effectively corrects abnormal skeleton points, flexibly embeds, and improves the accuracy of existing pose estimation algorithms. Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Peirong Xu, Wenzhong Guo |
AAAI | 4 |
| 2025 | URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image RestorationabstractExisting low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexible and effective degradation restoration for low-light images. Specifically, we customize the core URWKV block to perceive and analyze complex degradations by leveraging multiple intra- and inter-stage states. First, inspired by the pupil mechanism in the human visual system, we propose Luminance-adaptive Normalization (LAN) that adjusts normalization parameters based on rich inter-stage states, allowing for adaptive, scene-aware luminance modulation. Second, we aggregate multiple intra-stage states through exponential moving average approach, effectively capturing subtle variations while mitigating information loss inherent in the single-state mechanism. To reduce the degradation effects commonly associated with conventional skip connections, we propose the State-aware Selective Fusion (SSF) module, which dynamically aligns and integrates multi-state features across encoder stages, selectively fusing contextual information. In comparison to state-of-the-art models, our URWKV model achieves superior performance on various benchmarks, while requiring significantly fewer parameters and computational resources. Code is available at: https://github.com/FZU-N/URWKV. Rui Xu 0028, Yuzhen Niu, Yuezhou Li, Huangbiao Xu, Wenxi Liu, Yuzhong Chen 0001 |
CVPR | 1 |
| 2025 | Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentabstractLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large number of model parameters to learn potential associations between actions and music. To address this issue, we propose a language-guided audio-visual learning (MLAVL) framework that models "audio-action-visual" correlations guided by low-cost language modality. In our framework, multidimensional domain-based actions form action knowledge graphs, motivating audio-visual modalities to focus on task-relevant actions. We further design a shared-specific context encoder to integrate deep multimodal semantics, and an audio-visual cross-modal fusion module to evaluate action-music consistency. To match the sport’s rules, we then propose a dual-branch prompt-guided grading module to weigh both visual and audio-visual performance. Extensive experiments demonstrate that our approach achieves state-of-the-art on four public long-term sports benchmarks while maintaining low parameters.1 Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Wenzhong Guo |
CVPR | 4 |
| 2025 | IPCMoE: Integrating Perceptual Cues with Mixture-of-Experts for Joint Low-Light Image Enhancement and DeblurringabstractVisual perception of nighttime images is often compromised by co-existing low-light and blur degradations. While recent methods have made progress in jointly solving these degradations, the diversity of patterns and intensities in degradation has not been properly considered, leading to inconsistent illumination and unintended artifacts. In response, we propose to integrate perceptual cues with mixture-of-experts (IPCMoE) to achieve flexible processing for low-light blurry images. By exploiting the perceptual cues, we strategically combine dedicated experts with the selective collaboration approach for feature enlightening and texture restoration. To this end, we develop perceptual-integrated MoEs by designing customized routers and task-depended experts. Specifically, the texture memorial MoE is developed to preserve valuable features to restore high-fidelity details, and the enhancement MoE that adaptively integrates enlightening cues and texture cues is designed to formulate the relationship between feature enlightening and texture restoration, thereby achieving dynamic image processing. Extensive experiments show that our method achieves state-of-the-art performance on LOL-Blur and Real-LOL-Blur datasets. Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Hui Da, Rui Xu 0028, Wenxi Liu |
ACM Multimedia | 5 |
| 2025 | CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic Assessment
Yuzhen Niu, Siling Chen 0002, Yuzhong Chen 0001, Rui Xu 0028, Hui Da |
ACM Multimedia | 5 |
| 2025 | Parallax-aware dual-view feature enhancement and adaptive detail compensation for dual-pixel defocus deblurring
Yuzhen Niu, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | AMST: Object tracking based on collaborative framework with adaptive multi-strategy
Rui Xu 0028, Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
Inf. Sci. | 1 |
| 2025 | Collaboratively enhanced and integrated detail-context information for low-light image enhancement
Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Yuzhong Chen 0001 |
Pattern Recognit. | 4 |
| 2025 | Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual TrackingabstractIn recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods. Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Quality-Guided Vision-Language Learning for Long-Term Action Quality AssessmentabstractLong-term action quality assessment poses a challenging visual task since it requires assessing technical actions at different skill levels in a long video. Recent state-of-the-art methods incorporate additional modality information to aid in understanding action semantics, which incurs extra annotation costs and imposes higher constraints on action scenes and datasets. To address this issue, we propose a Quality-Guided Vision-Language Learning (QGVL) method to map visual features into appropriate fine-grained intervals of quality scores. Specifically, we use a set of quality-related textual prompts as quality prototypes to guide the discrimination and aggregation of specific visual actions. To avoid fuzzy rule mapping, we further propose a progressive semantic learning strategy with a Granularity-Adaptive Semantic Learning Module (GSLM) that refines accurate score intervals from coarse to fine at clip, grade, and score levels. The quality-related semantics we designed are universal to all types of action scenarios without any additional annotations. Extensive experiments show that our approach outperforms previous work by a significant margin and establishes new state-of-the-art on four public AQA benchmarks: Rhythmic Gymnastics, Fis-V, FS1000, and FineFS. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Yuezhou Li, Rui Xu 0028, Wenzhong Guo |
IEEE Trans. Multim. | 5 |
| 2024 | Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment
Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu 0028, Huanqi Wu 0001, Wenzhong Guo |
ECCV (42) | 4 |
| 2024 | MiNet: Weakly-Supervised Camouflaged Object Detection through Mutual Interaction between Region and Edge CuesabstractExisting weakly-supervised camouflaged object detection (WSCOD) methods have much difficulty in detecting accurate object boundaries due to insufficient and imprecise boundary supervision in scribble annotations. Drawing inspiration from human perception that discerns camouflaged objects by incorporating both object region and boundary information, we propose a novel Mutual Interaction Network (MiNet) for scribble-based WSCOD to alleviate the detection difficulty caused by insufficient scribbles. The proposed MiNet facilitates mutual reinforcement between region and edge cues, thereby integrating more robust priors to enhance detection accuracy. In this paper, we first construct an edge cue refinement net, featuring a core region-aware guidance module (RGM) aimed at leveraging the extracted region feature as a prior to generate the discriminative edge map. By considering both object semantic and positional relationships between edge feature and region feature, RGM highlights the areas associated with the object in the edge feature. Subsequently, to tackle the inherent similarity between camouflaged objects and the surroundings, we devise a region-boundary refinement net. This net incorporates a core edge-aware guidance module (EGM), which uses the enhanced edge map from the edge cue refinement net as guidance to refine the object boundaries in an iterative and multi-level manner. Experiments on CAMO, CHAMELEON, COD10K, and NC4K datasets demonstrate that the proposed MiNet outperforms the state-of-the-art methods. Yuzhen Niu, Lifen Yang, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001 |
ACM Multimedia | 3 |
| 2024 | Zero-Referenced Enlightening and Restoration for UAV Nighttime VisionabstractUnmanned aerial vehicle (UAV) based visual systems suffer from poor perception at nighttime. There are three challenges for enlightening nighttime vision for UAVs: Firstly, the UAV nighttime images differ from underexposed images in the statistical characteristic, limiting the performance of general low-light image enhancement (LLIE) methods. Secondly, when enlightening nighttime images, the artifacts tend to be amplified, distracting the visual perception of UAVs. Thirdly, due to the inherent scarcity of paired data in the real world, it is difficult for UAV nighttime vision to benefit from supervised learning. To meet these challenges, we propose a zero-referenced enlightening and restoration network (ZERNet) for improving the perception of UAV vision at nighttime. Specifically, by estimating the nighttime enlightening map (NE-map), a pixel-to-pixel transformation is then conducted to enlighten the dark pixels while suppressing overbright pixels. Furthermore, we propose the self-regularized restoration to preserve the semantic contents and restrict the artifacts in the final result. Finally, our method is derived from zero-referenced learning, which is free from paired training data. Comprehensive experiments show that the proposed ZERNet effectively improves the nighttime visual perception of UAVs on quantitative metrics, qualitative comparisons, and application-based analysis. Yuezhou Li, Yuzhen Niu, Rui Xu 0028 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | STD-Net: Spatio-Temporal Decomposition Network for Video Demoiréing With Sparse TransformersabstractThe problem of video demoiréing is a new challenge in video restoration. Unlike image demoiréing, which involves removing static and uniform patterns, video demoiréing requires tackling dynamic and varied moiré patterns while maintaining video details, colors, and temporal consistency. It is particularly challenging to model moiré patterns for videos with camera or object motions, where separating moiré from the original video content across frames is extremely difficult. Nonetheless, we observe that the spatial distribution of moiré patterns is often sparse on each frame, and their long-range temporal correlation is not significant. To fully leverage this phenomenon, a sparsity-constrained spatial self-attention scheme is proposed to concentrate on removing sparse moiré efficiently for each frame without being distracted by dynamic video content. The frame-wise spatial features are then correlated and aggregated via the local temporal cross-frame-attention module to produce temporal-consistent high-quality moiré-free videos. The above decoupled spatial and temporal transformers constitute the Spatio-Temporal Decomposition Network, dubbed STD-Net. For evaluation, we present a large-scale video demoiréing benchmark featuring various real-life scenes, camera motions, and object motions. We demonstrate that our proposed model can effectively and efficiently achieve superior performance on video demoiréing and single image demoiréing tasks.The proposed dataset will be released after the paper is accepted. Yuzhen Niu, Rui Xu 0028, Zhihua Lin, Wenxi Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Perceptual Decoupling With Heterogeneous Auxiliary Tasks for Joint Low-Light Image Enhancement and DeblurringabstractCapturing images at night are susceptible to inadequate illumination conditions and motion blurring. Given the typical coupling of these two forms of degradation, a pioneer work takes a compact approach of brightening followed by deblurring. However, this sequential approach may compromise informative features and elevate the likelihood of generating unintended artifacts. In this paper, we observe that the co-existing low light and blurs intuitively impair multiple perceptions, making it difficult to produce visually appealing results. To meet these challenges, we propose perceptual decoupling with heterogeneous auxiliary tasks (PDHAT) for joint low-light image enhancement and deblurring. Based on the crucial perceptual properties of the two degradations, we construct two individual auxiliary tasks: coarse preview prediction (CPP) and high-frequency reconstruction (HFR), so that the perception of color, brightness, edges, and details are decoupled into heterogeneous auxiliary tasks to obtain task-specific representations for parallel assisting the main task: joint low-light enhancement and deblurring (LLE-Deblur). Furthermore, we develop dedicated modules to build the network blocks in each branch based on the exclusive properties of each task. Comprehensive experiments are conducted on LOL-Blur and Real-LOL-Blur datasets, showing that our method outperforms existing methods on quantitative metrics and qualitative results. Yuezhou Li, Rui Xu 0028, Yuzhen Niu, Wenzhong Guo, Tiesong Zhao |
IEEE Trans. Multim. | 2 |
| 2024 | Bilateral Interaction for Local-Global Collaborative Perception in Low-Light Image EnhancementabstractLow-light image enhancement is a challenging task due to the limited visibility in dark environments. While recent advances have shown progress in integrating CNNs and Transformers, the inadequate local-global perceptual interactions still impedes their application in complex degradation scenarios. To tackle this issue, we propose BiFormer, a lightweight framework that facilitates local-global collaborative perception via bilateral interaction. Specifically, our framework introduces a core CNN-Transformer collaborative perception block (CPB) that combines local-aware convolutional attention (LCA) and global-aware recursive Transformer (GRT) to simultaneously preserve local details and ensure global consistency. To promote perceptual interaction, we adopt bilateral interaction strategy for both local and global perception, which involves local-to-global second-order interaction (SoI) in the dual-domain, as well as a mixed-channel fusion (MCF) module for global-to-local interaction. The MCF is also a highly efficient feature fusion module tailored for degraded features. Extensive experiments conducted on low-level and high-level tasks demonstrate that BiFormer achieves state-of-the-art performance. Furthermore, it exhibits a significant reduction in model parameters and computational cost compared to existing Transformer-based low-light image enhancement methods. Rui Xu 0028, Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Yuzhong Chen 0001, Tiesong Zhao |
IEEE Trans. Multim. | 1 |
| 2023 | Zero-referenced low-light image enhancement with adaptive filter network
Yuezhou Li, Yuzhen Niu, Rui Xu 0028, Yuzhong Chen 0001 |
Eng. Appl. Artif. Intell. | 3 |