VLDB 2026 Research / reviewers in the wild / expert
Yuezhou Li
dblp:284/2562
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-7397-4661ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flare detection and detail compensation for nighttime flare removal
Yuzhen Niu, Yuezhou Li, Jingyuan Zheng |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Integrating perceptual cues with mixture-of-experts for low-light image restoration
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Hui Da, Wenxi Liu, Lifang Wei |
Neural Networks | 1 |
| 2025 | DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseabstractThe fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix proposes a bidirectional spatial-temporal context optical flow correction (BOFC) method. This approach leverages the consistency and complementarity of motion information between two modalities: optical flow, which excels at pixel capture, and lightweight skeleton data. It enables the extraction of pixel-level motion changes and the correction of abnormal skeleton data. Furthermore, we propose a part-level dance dataset (Dancer Parts) and part-level motion feature extraction based on task decoupling (PETD). This aims to decouple complex whole-body parts tracking into fine-grained limb-level motion extraction, enhancing the confidence of temporal information and the accuracy of correction for abnormal data. Finally, we present the DNV dataset, which simulates fully neat group dance scenes and provides reliable labels and validation methods for the newly introduced group dance neatness assessment (GDNA). To the best of our knowledge, this is the first work to develop quantitative criteria for assessing limb and joint neatness in group dance. We conduct experiments on DNV and video-based public JHMDB datasets. Our method effectively corrects abnormal skeleton points, flexibly embeds, and improves the accuracy of existing pose estimation algorithms. Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Peirong Xu, Wenzhong Guo |
AAAI | 5 |
| 2025 | URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image RestorationabstractExisting low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexible and effective degradation restoration for low-light images. Specifically, we customize the core URWKV block to perceive and analyze complex degradations by leveraging multiple intra- and inter-stage states. First, inspired by the pupil mechanism in the human visual system, we propose Luminance-adaptive Normalization (LAN) that adjusts normalization parameters based on rich inter-stage states, allowing for adaptive, scene-aware luminance modulation. Second, we aggregate multiple intra-stage states through exponential moving average approach, effectively capturing subtle variations while mitigating information loss inherent in the single-state mechanism. To reduce the degradation effects commonly associated with conventional skip connections, we propose the State-aware Selective Fusion (SSF) module, which dynamically aligns and integrates multi-state features across encoder stages, selectively fusing contextual information. In comparison to state-of-the-art models, our URWKV model achieves superior performance on various benchmarks, while requiring significantly fewer parameters and computational resources. Code is available at: https://github.com/FZU-N/URWKV. Rui Xu 0028, Yuzhen Niu, Yuezhou Li, Huangbiao Xu, Wenxi Liu, Yuzhong Chen 0001 |
CVPR | 3 |
| 2025 | Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentabstractLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large number of model parameters to learn potential associations between actions and music. To address this issue, we propose a language-guided audio-visual learning (MLAVL) framework that models "audio-action-visual" correlations guided by low-cost language modality. In our framework, multidimensional domain-based actions form action knowledge graphs, motivating audio-visual modalities to focus on task-relevant actions. We further design a shared-specific context encoder to integrate deep multimodal semantics, and an audio-visual cross-modal fusion module to evaluate action-music consistency. To match the sport’s rules, we then propose a dual-branch prompt-guided grading module to weigh both visual and audio-visual performance. Extensive experiments demonstrate that our approach achieves state-of-the-art on four public long-term sports benchmarks while maintaining low parameters.1 Huangbiao Xu, Xiao Ke, Huanqi Wu 0001, Rui Xu 0028, Yuezhou Li, Wenzhong Guo |
CVPR | 5 |
| 2025 | IPCMoE: Integrating Perceptual Cues with Mixture-of-Experts for Joint Low-Light Image Enhancement and DeblurringabstractVisual perception of nighttime images is often compromised by co-existing low-light and blur degradations. While recent methods have made progress in jointly solving these degradations, the diversity of patterns and intensities in degradation has not been properly considered, leading to inconsistent illumination and unintended artifacts. In response, we propose to integrate perceptual cues with mixture-of-experts (IPCMoE) to achieve flexible processing for low-light blurry images. By exploiting the perceptual cues, we strategically combine dedicated experts with the selective collaboration approach for feature enlightening and texture restoration. To this end, we develop perceptual-integrated MoEs by designing customized routers and task-depended experts. Specifically, the texture memorial MoE is developed to preserve valuable features to restore high-fidelity details, and the enhancement MoE that adaptively integrates enlightening cues and texture cues is designed to formulate the relationship between feature enlightening and texture restoration, thereby achieving dynamic image processing. Extensive experiments show that our method achieves state-of-the-art performance on LOL-Blur and Real-LOL-Blur datasets. Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Hui Da, Rui Xu 0028, Wenxi Liu |
ACM Multimedia | 1 |
| 2025 | Parallax-aware dual-view feature enhancement and adaptive detail compensation for dual-pixel defocus deblurring
Yuzhen Niu, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Quality-Guided Vision-Language Learning for Long-Term Action Quality AssessmentabstractLong-term action quality assessment poses a challenging visual task since it requires assessing technical actions at different skill levels in a long video. Recent state-of-the-art methods incorporate additional modality information to aid in understanding action semantics, which incurs extra annotation costs and imposes higher constraints on action scenes and datasets. To address this issue, we propose a Quality-Guided Vision-Language Learning (QGVL) method to map visual features into appropriate fine-grained intervals of quality scores. Specifically, we use a set of quality-related textual prompts as quality prototypes to guide the discrimination and aggregation of specific visual actions. To avoid fuzzy rule mapping, we further propose a progressive semantic learning strategy with a Granularity-Adaptive Semantic Learning Module (GSLM) that refines accurate score intervals from coarse to fine at clip, grade, and score levels. The quality-related semantics we designed are universal to all types of action scenarios without any additional annotations. Extensive experiments show that our approach outperforms previous work by a significant margin and establishes new state-of-the-art on four public AQA benchmarks: Rhythmic Gymnastics, Fis-V, FS1000, and FineFS. Huangbiao Xu, Huanqi Wu 0001, Xiao Ke, Yuezhou Li, Rui Xu 0028, Wenzhong Guo |
IEEE Trans. Multim. | 4 |
| 2025 | Skeleton-Boundary-Guided Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to resolve the tough issue of accurately segmenting objects hidden in the surroundings. However, the existing methods suffer from two major problems: the incomplete interior and the inaccurate boundary of the object. To address these difficulties, we propose a three-stage skeleton-boundary–guided network (SBGNet) for the COD task. Specifically, we design a novel skeleton-boundary label to be complementary to the typical pixel-wise mask annotation, emphasizing the interior skeleton and the boundary of the camouflaged object. Furthermore, the proposed feature guidance module (FGM) leverages the skeleton-boundary feature to guide the model to focus on both the interior and the boundary of the camouflaged object. Besides, we design a bidirectional feature flow path with the information interaction module (IIM) to propagate and integrate the semantic and texture information. Finally, we propose the dual feature distillation module (DFDM) to progressively refine the segmentation results in a fine-grained manner. Comprehensive experiments demonstrate that our SBGNet outperforms 20 state-of-the-art methods on three benchmarks in both qualitative and quantitative comparisons. Yuzhen Niu, Yeyuan Xu, Yuezhou Li, Jiabang Zhang, Yuzhong Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment
Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu 0028, Huanqi Wu 0001, Wenzhong Guo |
ECCV (42) | 3 |
| 2024 | MiNet: Weakly-Supervised Camouflaged Object Detection through Mutual Interaction between Region and Edge CuesabstractExisting weakly-supervised camouflaged object detection (WSCOD) methods have much difficulty in detecting accurate object boundaries due to insufficient and imprecise boundary supervision in scribble annotations. Drawing inspiration from human perception that discerns camouflaged objects by incorporating both object region and boundary information, we propose a novel Mutual Interaction Network (MiNet) for scribble-based WSCOD to alleviate the detection difficulty caused by insufficient scribbles. The proposed MiNet facilitates mutual reinforcement between region and edge cues, thereby integrating more robust priors to enhance detection accuracy. In this paper, we first construct an edge cue refinement net, featuring a core region-aware guidance module (RGM) aimed at leveraging the extracted region feature as a prior to generate the discriminative edge map. By considering both object semantic and positional relationships between edge feature and region feature, RGM highlights the areas associated with the object in the edge feature. Subsequently, to tackle the inherent similarity between camouflaged objects and the surroundings, we devise a region-boundary refinement net. This net incorporates a core edge-aware guidance module (EGM), which uses the enhanced edge map from the edge cue refinement net as guidance to refine the object boundaries in an iterative and multi-level manner. Experiments on CAMO, CHAMELEON, COD10K, and NC4K datasets demonstrate that the proposed MiNet outperforms the state-of-the-art methods. Yuzhen Niu, Lifen Yang, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001 |
ACM Multimedia | 4 |
| 2024 | Zero-Referenced Enlightening and Restoration for UAV Nighttime VisionabstractUnmanned aerial vehicle (UAV) based visual systems suffer from poor perception at nighttime. There are three challenges for enlightening nighttime vision for UAVs: Firstly, the UAV nighttime images differ from underexposed images in the statistical characteristic, limiting the performance of general low-light image enhancement (LLIE) methods. Secondly, when enlightening nighttime images, the artifacts tend to be amplified, distracting the visual perception of UAVs. Thirdly, due to the inherent scarcity of paired data in the real world, it is difficult for UAV nighttime vision to benefit from supervised learning. To meet these challenges, we propose a zero-referenced enlightening and restoration network (ZERNet) for improving the perception of UAV vision at nighttime. Specifically, by estimating the nighttime enlightening map (NE-map), a pixel-to-pixel transformation is then conducted to enlighten the dark pixels while suppressing overbright pixels. Furthermore, we propose the self-regularized restoration to preserve the semantic contents and restrict the artifacts in the final result. Finally, our method is derived from zero-referenced learning, which is free from paired training data. Comprehensive experiments show that the proposed ZERNet effectively improves the nighttime visual perception of UAVs on quantitative metrics, qualitative comparisons, and application-based analysis. Yuezhou Li, Yuzhen Niu, Rui Xu 0028 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Perceptual Decoupling With Heterogeneous Auxiliary Tasks for Joint Low-Light Image Enhancement and DeblurringabstractCapturing images at night are susceptible to inadequate illumination conditions and motion blurring. Given the typical coupling of these two forms of degradation, a pioneer work takes a compact approach of brightening followed by deblurring. However, this sequential approach may compromise informative features and elevate the likelihood of generating unintended artifacts. In this paper, we observe that the co-existing low light and blurs intuitively impair multiple perceptions, making it difficult to produce visually appealing results. To meet these challenges, we propose perceptual decoupling with heterogeneous auxiliary tasks (PDHAT) for joint low-light image enhancement and deblurring. Based on the crucial perceptual properties of the two degradations, we construct two individual auxiliary tasks: coarse preview prediction (CPP) and high-frequency reconstruction (HFR), so that the perception of color, brightness, edges, and details are decoupled into heterogeneous auxiliary tasks to obtain task-specific representations for parallel assisting the main task: joint low-light enhancement and deblurring (LLE-Deblur). Furthermore, we develop dedicated modules to build the network blocks in each branch based on the exclusive properties of each task. Comprehensive experiments are conducted on LOL-Blur and Real-LOL-Blur datasets, showing that our method outperforms existing methods on quantitative metrics and qualitative results. Yuezhou Li, Rui Xu 0028, Yuzhen Niu, Wenzhong Guo, Tiesong Zhao |
IEEE Trans. Multim. | 1 |
| 2024 | Bilateral Interaction for Local-Global Collaborative Perception in Low-Light Image EnhancementabstractLow-light image enhancement is a challenging task due to the limited visibility in dark environments. While recent advances have shown progress in integrating CNNs and Transformers, the inadequate local-global perceptual interactions still impedes their application in complex degradation scenarios. To tackle this issue, we propose BiFormer, a lightweight framework that facilitates local-global collaborative perception via bilateral interaction. Specifically, our framework introduces a core CNN-Transformer collaborative perception block (CPB) that combines local-aware convolutional attention (LCA) and global-aware recursive Transformer (GRT) to simultaneously preserve local details and ensure global consistency. To promote perceptual interaction, we adopt bilateral interaction strategy for both local and global perception, which involves local-to-global second-order interaction (SoI) in the dual-domain, as well as a mixed-channel fusion (MCF) module for global-to-local interaction. The MCF is also a highly efficient feature fusion module tailored for degraded features. Extensive experiments conducted on low-level and high-level tasks demonstrate that BiFormer achieves state-of-the-art performance. Furthermore, it exhibits a significant reduction in model parameters and computational cost compared to existing Transformer-based low-light image enhancement methods. Rui Xu 0028, Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Yuzhong Chen 0001, Tiesong Zhao |
IEEE Trans. Multim. | 2 |
| 2023 | Zero-referenced low-light image enhancement with adaptive filter network
Yuezhou Li, Yuzhen Niu, Rui Xu 0028, Yuzhong Chen 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Efficient Encoder-Decoder Network With Estimated Direction for SAR Ship DetectionabstractSynthetic aperture radar (SAR) image ship detection has important applications in marine surveillance. There are two limitations when applying advanced detection methods naively for SAR ship detection. First, most detectors construct the model as an encoder and rely on the feature pyramid network (FPN) head for accurate prediction, which may lead to high computational costs. Second, the background noises in the ground truth (annotated as rectangular bounding boxes) of angular ships bring difficulties for model training. To meet these challenges, we propose an efficient encoder–decoder network with estimated direction for ship detection in SAR images. First, we present an anchor-free encoder–decoder model that can efficiently extract multiple-level features. Second, we formulate ship detection as a multitask learning problem, including a bounding box prediction and a ship direction regression. The estimated ship direction can weakly supervise and benefit ship detection. Furthermore, we develop a center-weighted labeling method for overlapped annotations. Comprehensive experiments on SAR-Ship-Detection and SSDD datasets show that our method achieves state-of-the-art performance with a high running speed. Yuzhen Niu, Yuezhou Li, Jiangyi Huang, Yuzhong Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Learning deep convolutional descriptor aggregation for efficient visual tracking
Xiao Ke, Yuezhou Li, Wenzhong Guo, Yanyan Huang |
Neural Comput. Appl. | 2 |
| 2021 | Template Enhancement and Mask Generation for Siamese TrackingabstractSiamese tracking methods have become the focus of visual tracking in recent years. Advanced Siamese trackers perform well on certain benchmarks, but there are still some limitations. First, most Siamese trackers adopt the initial frame as a single template, which leads to underfitting and reduces the ability to predict instances. Second, mainstream trackers report a rectangular bounding box as a prediction, resulting in poor accuracy of non-rigid objects. Therefore, we propose the template enhancement and mask generation for Siamese tracking. Given that the essence of Siamese trackers is instance learning, we propose constructing an alternative template explicitly to address the underfitting of the instance space. Moreover, in order to improve the tracking accuracy, we obtain the descriptor aggregation to transform the semantic segmentation outputs for mask prediction. Finally, we propose the SiamEM through the fusion of the above approaches. Comprehensive experiments show that template enhancement and mask generation significantly improve Siamese trackers on benchmarks. Xiao Ke, Yuezhou Li, Yu Ye 0004, Wenzhong Guo |
IEEE Signal Process. Lett. | 2 |