EDBT 2026 Demo / reviewers in the wild / expert
Wu Ran
dblp:272/6751
· DBLP profile ↗
20ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-8478-0750ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Align-then-generate: An effective cross-modal generation paradigm for multi-label zero-shot learning
Peirong Ma, Wu Ran, Yanhui Gu, Huaqiu Chen, Zhiquan He, Hong Lu 0001 |
Pattern Recognit. | 2 |
| 2025 | Cross-Architecture Distillation Made Simple with Redundancy Suppression
Yuehao Liu, Wu Ran, Chao Ma 0004 |
ICCV | 3 |
| 2025 | Implicit Retinex Decomposition with Chromaticity Disentanglement for Low-Light Image Enhancement
Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma |
ACM Multimedia | 2 |
| 2025 | Correlated Low-Rank Adaptation for ConvNetsabstractLow-Rank Adaptation (LoRA) methods have demonstrated considerable success in achieving parameter-efficient fine-tuning (PEFT) for Transformer-based foundation models. These methods typically fine-tune individual Transformer layers using independent LoRA adaptations. However, directly applying existing LoRA techniques to convolutional networks (ConvNets) yields unsatisfactory results due to the high correlation between the stacked sequential layers of ConvNets. To overcome this challenge, we introduce a novel framework called Correlated Low-Rank Adaptation (CoLoRA), which explicitly utilizes correlated low-rank matrices to model the inter-layer dependencies among convolutional layers. Additionally, to enhance tuning efficiency, we propose a parameter-free filtering method that enlarges the receptive field of LoRA, thus minimizing interference from non-informative local regions. Comprehensive experiments conducted across various mainstream vision tasks, including image classification, semantic segmentation, and object detection, illustrate that CoLoRA significantly advances the state-of-the-art PEFT approaches. Notably, our CoLoRA achieves superior performance with only 5\% of trainable parameters, surpassing full fine-tuning in the image classification task on the VTAB-1k dataset using ConvNeXt-S. Code is available at [https://github.com/VISION-SJTU/CoLoRA](https://github.com/VISION-SJTU/CoLoRA). Wu Ran, Shuyang Pang, Jinfan Liu, Jingsheng Liu, Yichao Yan, Chao Ma 0004 |
NeurIPS | 1 |
| 2025 | Unleashing the Potential of Hierarchical Region Clues for Open-Vocabulary Multi-Label ClassificationabstractOpen-vocabulary multi-label classification (OV-MLC) aims to leverage the rich multi-modal knowledge from Vision-language pre-training (VLP) models to further improve the recognition ability for unseen (novel) classes beyond the training set in multi-label scenarios. Existing OV-MLC methods only perform predictions on single hierarchical regions, and aggregate the prediction scores of these regions through simpletop-kmean pooling. This fails to unleash the potential of rich hierarchical region clues in multi-label images and does not fully exploit the discriminative information from all regions in the image, resulting in sub-optimal performance. In this work, we propose a novel OV-MLC framework to fully harness the power of multiple hierarchical region clues. Specifically, we first design a hierarchical clue gathering (HCG) module to gather different hierarchical clues, enabling more precise recognition of multiple object categories with different sizes in a multi-label image. Then, by viewing multi-label classification as single-label classification of each region within the image, we present a novel hierarchical score aggregation (HSA) approach, thereby better utilizing the predictions of each image region for each class. We also utilize a well-designed region selection strategy (RSS) to eliminate noise or background regions in an image that are irrelevant to classification, achieving higher multi-label classification accuracy. In addition, we propose a hybrid prompt learning (HPL) strategy to enhance visual-semantic consistency while preserving the generalization capability of label embeddings for unseen classes. Extensive experiments on public benchmark datasets demonstrate that our method significantly outperforms the current state-of-the-art. Peirong Ma, Wu Ran, Zhiquan He, Jian Pu, Hong Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate RainsabstractRecent advances in image deraining have focused on training powerful models on mixed multiple datasets comprising diverse rain types and backgrounds. However, this approach tends to overlook the inherent differences among rainy images, leading to suboptimal results. To overcome this limitation, we focus on addressing various rainy images by delving into meaningful representations that encapsulate both the rain and background components. Leveraging these representations as instructive guidance, we put forth a Context-based Instance-level Modulation (CoI-M) mechanism adept at efficiently modulating CNN- or Transformer-based models. Furthermore, we devise a rain-/detail-aware contrastive learning strategy to help extract joint rain-/detail-aware representations. By integrating CoI-M with the rain-/detail-aware Contrastive learning, we develop [CoIC](https://github.com/Schizophreni/CoIC), an innovative and potent algorithm tailored for training models on mixed datasets. Moreover, CoIC offers insight into modeling relationships of datasets, quantitatively assessing the impact of rain and details on restoration, and unveiling distinct behaviors of models given diverse inputs. Extensive experiments validate the efficacy of CoIC in boosting the deraining ability of CNN and Transformer models. CoIC also enhances the deraining prowess remarkably when real-world dataset is included. Wu Ran, Peirong Ma, Zhiquan He, Hao Ren 0002, Hong Lu 0001 |
ICLR | 1 |
| 2024 | Rainmer: Learning Multi-view Representations for Comprehensive Image Deraining and BeyondabstractWe address image deraining under complex backgrounds, diverse rain scenarios, and varying illumination conditions, representing a highly practical and challenging problem. Our approach utilizes synthetic, real-world, and nighttime datasets, wherein rich backgrounds, multiple degradation types, and diverse illumination conditions coexist. The primary challenge in training models on these datasets arises from the discrepancies among them, potentially leading to conflicts or competition during the training period. To address this issue, we first align the distribution of synthetic, real-world and nighttime datasets. Then we propose a novel contrastive learning strategy to extract multi-view (multiple) representations that effectively capture image details, degradations, and illuminations, thereby facilitating training across all datasets. Regarding multiple representations as profitable prompts for deraining, we devise a prompting strategy to integrate them into the decoding process. This contributes to a potent deraining model, dubbed Rainmer. Additionally, a spatial-channel interaction module is introduced to fully exploit cues when extracting multi-view representations. Extensive experiments on synthetic, real-world, and nighttime datasets demonstrate that Rainmer outperforms current representative methods. Moreover, Rainmer achieves superior performance on the All-in-One image restoration dataset, underscoring its effectiveness. Furthermore, quantitative results reveal that Rainmer significantly improves object detection performance on both daytime and nighttime rainy datasets. These observations substantiate the potential of Rainmer for practical applications. Wu Ran, Peirong Ma, Zhiquan He, Hong Lu 0001 |
ACM Multimedia | 1 |
| 2024 | Feature decoupling and reorganization network for single image deraining
Yunrui Cheng, Junjian Huang, Hao Ren 0002, Wu Ran, Hong Lu 0001 |
Multim. Syst. | 4 |
| 2024 | Low-Light Image Enhancement With Multi-Scale Attention and Frequency-Domain OptimizationabstractLow-light image enhancement aims to improve the perceptual quality of images captured in conditions of insufficient illumination. However, such images are often characterized by low visibility and noise, making the task challenging. Recently, significant progress has been made using deep learning-based approaches. Nonetheless, existing methods encounter difficulties in balancing global and local illumination enhancement and may fail to suppress noise in complex lighting conditions. To address these issues, we first propose a multi-scale illumination adjustment network to balance both global illumination and local contrast. Furthermore, to effectively suppress noise potentially amplified by the illumination adjustment, we introduce a wavelet-based attention network that efficiently perceives and removes noise in the frequency domain. We additionally incorporate a discrete wavelet transform loss to supervise the training process. Particularly, the proposed wavelet-based attention network has been shown to enhance the performance of existing low-light image enhancement methods. This observation indicates that the proposed wavelet-based attention network can be flexibly adapted to current approaches to yield superior enhancement results. Furthermore, extensive experiments conducted on benchmark datasets and downstream object detection task demonstrate that our proposed method achieves state-of-the-art performance and generalization ability. Zhiquan He, Wu Ran, Kehua Li, Chang-Yong Xie, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A Transferable Generative Framework for Multi-Label Zero-Shot LearningabstractMulti-label zero-shot learning (MLZSL) is a more realistic and challenging task than single-label zero-shot learning (SLZSL), which aims to recognize multiple unseen classes in a single image. To adapt generative models to the MLZSL task and better recognize multiple unseen object categories in an image, this paper proposes a Transferable Generative Framework (TGF), which consists of a Multi-Label Semantic Embedding Autoencoders (SEAs), a Semantic-Related Multi-Label Feature Transformation Network (FTN) and a Multi-Label Feature Generation Networks (FGNs). First, SEAs adaptively encodes the class-level word vectors corresponding to each sample containing different number of classes into sample-level semantic embeddings with the same dimension. Then, FTN transforms global features extracted by a CNN pre-trained on single-label images into features that are semantic-related and more suitable for multi-label classification. Finally, FGNs generates both global and local features to better recognize the dominant and minor object categories in a multi-label image, respectively. Extensive experiments on three benchmark datasets show that TGF significantly outperforms state-of-the-arts. Specifically, compared with the previous best generative MLZSL method (i.e., Gen-MLZSL), TGF improves the mAP of the ZSL (GZSL) task by 5.4% (6.9%), 20.5% (27.9%), and 2.4% (3.9%) on NUS-WIDE, Open Images, and MS-COCO datasets, respectively. Peirong Ma, Zhiquan He, Wu Ran, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Weakly-supervised Temporal Action Localization with Adaptive Clustering and Refining NetworkabstractWeakly-supervised temporal action localization task aims to localize temporal boundaries of action instances by using only video-level labels. Existing methods primarily adopt Multi-Instance-Learning (MIL) scheme to handle this task. The effectiveness of MIL scheme depends heavily on the selection of top-k action snippets, which is unstable and requires manual tuning. To address these deficiencies, we propose an Adaptive Clustering and Refining Network (ACRNet). Specifically, we present an action-aware clustering strategy that is adaptable and requires no manual tuning to separate action and background snippets of diverse videos based on intra-class activation distribution. And a cluster refining step is included to eliminate false action snippets by considering inter-class activation distribution, which greatly improves robustness and localization accuracy. Extensive experiments on THUMOS14, ActivityNet 1.2&1.3 benchmarks show that our method achieves state-of-the-art performance. Hao Ren 0002, Wu Ran, Xingson Liu, Hong Lu 0001, Cheng Jin 0001 |
ICME | 2 |
| 2023 | TRNR: Task-Driven Image Rain and Noise Removal With a Few Images Based on Patch AnalysisabstractThe recent success of learning-based image rain and noise removal can be attributed primarily to well-designed neural network architectures and large labeled datasets. However, we discover that current image rain and noise removal methods result in low utilization of images. To alleviate the reliance of deep models on large labeled datasets, we propose the task-driven image rain and noise removal (TRNR) based on a patch analysis strategy. The patch analysis strategy samples image patches with various spatial and statistical properties for training and can increase image utilization. Furthermore, the patch analysis strategy encourages us to introduce the N-frequency-K-shot learning task for the task-driven approach TRNR. TRNR allows neural networks to learn from numerous N-frequency-K-shot learning tasks, rather than from a large amount of data. To verify the effectiveness of TRNR, we build a Multi-Scale Residual Network (MSResNet) for both image rain removal and Gaussian noise removal. Specifically, we train MSResNet for image rain removal and noise removal with a few images (for example, 20.0% train-set of Rain100H). Experimental results demonstrate that TRNR enables MSResNet to learn more effectively when data is scarce. TRNR has also been shown in experiments to improve the performance of existing methods. Furthermore, MSResNet trained with a few images using TRNR outperforms most recent deep learning methods trained data-driven on large labeled datasets. These experimental results have confirmed the effectiveness and superiority of the proposed TRNR. The source code is available on https://github.com/Schizophreni/MSResNet-TRNR. Wu Ran, Bohong Yang, Peirong Ma, Hong Lu 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Semantic-Related Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) is a challenging task. Although no visual samples of unseen classes are provided during training, the classifier must learn to recognize all classes (i.e. both seen and unseen classes). Due to the ability to generate unseen classes samples, generative models have been widely used in GZSL. However, these generative models only learn from the seen classes, so the discriminability of the unseen class features they generate is usually poor, resulting in low unseen class classification accuracy. To solve this problem, this paper proposes a novel semantic-related feature generative (SRFG) model to improve visual-semantic consistency and alleviate seen-unseen bias effectively. SRFG can generate any number of semantic-related discriminative features for both seen and unseen classes. Extensive experiments on four benchmark datasets show that the proposed model significantly outperforms the state of the arts. Peirong Ma, Wu Ran, Hong Lu 0001 |
ICME | 2 |
| 2022 | Hybrid Uncalibrated Near-light Photometric Stereo in Realistic EnvironmentabstractPhotometric stereo aims at recovering the surface and shape of an object from a set of observations under different light conditions. Deep learning based methods have made substantial contributions to the surface normal estimation under complex surface reflections, various materials, near-light settings, etc. These deep learning based methods learn the mapping from observed images to the surface normal of an object directly via training deep neural networks on large labeled datasets. However, the shapes, materials, and reflectance properties of objects in these datasets are limited, leading to abridged performances in realistic environments. In this paper, we introduce a work-piece dataset for near-light photometric stereo under industrial application scenarios, which consists of observed images taken under at most 40 light conditions, and the ground truth surface normals. Based on this datasets, we propose the Hybrid Near-light Uncalibrated Photometric Stereo (HNUPS) for both unsupervised light calibration and surface normal estimation. Experimental results on the work-piece dataset demonstrate that HNUPS can obtain the least mean angular error when compared to recent photometric stereo methods, which have verified the effectiveness of the proposed HNUPS. Wu Ran, Xingsong Liu, Hong Lu 0001, Bohong Yang, Jingjing Luo |
ISCAS | 1 |
| 2022 | Weakly-Supervised Temporal Action Localization with Multi-Head Cross-Modal Attention
Hao Ren 0002, Wu Ran, Hong Lu 0001, Cheng Jin 0001 |
PRICAI (3) | 3 |
| 2022 | GAN-MVAE: A discriminative latent feature generation framework for generalized zero-shot learning
Peirong Ma, Hong Lu 0001, Bohong Yang, Wu Ran |
Pattern Recognit. Lett. | 4 |
| 2022 | Multi-Classes and Motion Properties for Concurrent Visual SLAM in Dynamic EnvironmentsabstractWorking in a dynamic environment is a challenging problem for visual simultaneous localization and mapping (visual SLAM). Most of the existing visual SLAM algorithms fail resulting in significant error or losing in tracking when moving objects dominate the scene. We found two reasons cause these issues: (i) Previous approaches use information from all regions in the image; (ii) Existing algorithms use just two groups and block all feature points from moveable objects. In this paper, we propose a novel Multi-classes and motion properties for Concurrent Visual SLAM (MCV-SLAM) algorithm, which defines classes into five categories and concurrently fuses prior knowledge and observation of moving objects with semantic segmentation to ensure visual SLAM works properly for dynamic environments in real time. We also propose an adaptive method to optimize camera pose by using more potential inlier feature points with continuous weights, while eliminating the impact of moving objects. Our experiments are performed on public datasets of both indoor and outdoor scenes with moving objects in dynamic environments. The experimental results demonstrate that our method outperforms previous works with greater robustness and smaller tracking errors, and our MCV-SLAM can deal with the situations (i.e., the dominance of moving objects, lack of matching points), which lead misestimating occurs in existing SLAMs. Bohong Yang, Wu Ran, Lin Wang 0033, Hong Lu 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Multim. | 2 |
| 2021 | Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action RecognitionabstractRecent attempts show that factorizing 3D convolutional filters into separate spatial and temporal components brings impressive improvement in action recognition. However, traditional temporal convolution operating along the temporal dimension will aggregate unrelated features, since the feature maps of fast-moving objects have shifted spatial positions. In this paper, we propose a novel and effective Multi-Directional Convolution (MDConv), which extracts features along different spatial-temporal orientations. Especially, MDConv has the same FLOPs and parameters as the traditional 1D temporal convolution. Also, we propose the Spatial-Temporal Feature Pyramid Module (STFPM) to fuse spatial semantics in different scales in a light-weight way. Our extensive experiments show that the models which integrate with MDConv achieve better accuracy on several large-scale action recognition benchmarks such as Kinetics, AVA and Something-Something V1&V2 datasets. Bohong Yang, Wu Ran, Hong Lu 0001, Yi-Ping Phoebe Chen |
ICASSP | 3 |
| 2020 | Single Image Rain Removal Boosting Via Directional GradientabstractImage rain removal has been widely studied with traditional methods and learning based methods for years. However, traditional methods like Gaussian mixture model and dictionary learning methods are time consuming and fail to well tackle images with heavy rain streaks since image patches are severely contaminated. By considering the line-like property and angle distribution of rain streaks, this problem can be well solved. In this paper, by introducing Directional b radient operator of arbitrary direction, we propose an efficient and robust Constraints based Model (DiG-CoM) for single image rain removal. Moreover, a density metric of rain streaks is applied to generalize the proposed model to light and heavy rain streak occasions. Extensive experiments on synthetic datasets demonstrate that the proposed model outperforms GMM and JCAS while requiring less time. Furthermore, on real-world occasions, the proposed method obtains better generalization ability compared with the stateof-the-art learning based methods. The source code is available at https://github.com/Schizophreni/Set-vanish-to-the-rain. Wu Ran, Youzhao Yang, Hong Lu 0001 |
ICME | 1 |
| 2020 | Rddan: A Residual Dense Dilated Aggregated Network For Single Image DerainingabstractRainy images contain rain streaks with different sizes, shapes, directions, and densities. To efficiently remove rain streaks from rainy images, it is necessary to capture rich rain details. In this paper, we propose a Residual Dense Dilated Aggregated Network (RDDAN) to focus on different types of rain steaks and efficiently model rain distribution from rainy images. Specifically, a Residual Dense Dilated Aggregated Block (RDDAB) is constructed to fully extract and exploit rain details hierarchically. In RDDAB, dilated aggregated module is applied to capture multi-scale rain details, dense connection is employed to fully exploit hierarchical features extracted by dilated aggregated module, and residual connection is introduced to keep flow of rain details among different blocks. Besides, all the features extracted by each RDDAB are fused progressively which allows the network to adaptively focus on significant hierarchical features inter blocks. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on synthetic and real-world datasets. The source code is available at https://github.com/nnUyi/RDDAN. Youzhao Yang, Wu Ran, Hong Lu 0001 |
ICME | 2 |