Hao Ren 0002

dblp:02/5283-2 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-5639-450XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Low-Light Image Enhancement via Multi-Exposure Progressive Contrastive Regularization
abstract
Low-light image enhancement (LLIE) aims to restore low-light images to their normal-light counterparts with optimal global illumination distribution and clear local details. With the advancement of deep learning, deep learning-based methods have become the mainstream in the LLIE community. However, most deep learning-based method cannot yet fully exploit the global and local contextual information in the low-light image. In this paper, we introduce a dual-branch module to simultaneously restore global and local features from spatial and frequency domain. To fuse these multi-level features, we propose a perception module to perform feature interaction between global and local features via cross attention and self-gating. By integrating the two developed modules into a U-Net backbone, we present a global-local interaction network for LLIE. Furthermore, recent studies have shown that contrastive learning can be an effective paradigm for the LLIE task. However, previous works typically use semantically-inconsistent under-/over-exposed images as negative samples. These images are very dissimilar to the ground-truth and cannot provide sufficient regularization in contrastive learning. To address this limitation, we explore a practical multi-exposure progressive contrastive regularization framework for LLIE. With a customized sample generation, sample selection, and progressive learning strategy, our proposed framework progressively narrows down the solution space around the optimum, and helps to improve the performance of LLIE methods without additional inference overhead. Combining the proposed network and contrastive regularization, our proposed method achieves favorable results compared to state-of-the-art LLIE methods on benchmark datasets. Extensive experiments further demonstrate the generalization ability of our proposed method.
Zuojie Xie, Hao Ren 0002, Junjian Huang, Zhiquan He, Hong Lu 0001, Lvfan Yuan, Changyong Xie
IEEE Trans. Circuits Syst. Video Technol.2
2024 Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate Rains
abstract
Recent advances in image deraining have focused on training powerful models on mixed multiple datasets comprising diverse rain types and backgrounds. However, this approach tends to overlook the inherent differences among rainy images, leading to suboptimal results. To overcome this limitation, we focus on addressing various rainy images by delving into meaningful representations that encapsulate both the rain and background components. Leveraging these representations as instructive guidance, we put forth a Context-based Instance-level Modulation (CoI-M) mechanism adept at efficiently modulating CNN- or Transformer-based models. Furthermore, we devise a rain-/detail-aware contrastive learning strategy to help extract joint rain-/detail-aware representations. By integrating CoI-M with the rain-/detail-aware Contrastive learning, we develop [CoIC](https://github.com/Schizophreni/CoIC), an innovative and potent algorithm tailored for training models on mixed datasets. Moreover, CoIC offers insight into modeling relationships of datasets, quantitatively assessing the impact of rain and details on restoration, and unveiling distinct behaviors of models given diverse inputs. Extensive experiments validate the efficacy of CoIC in boosting the deraining ability of CNN and Transformer models. CoIC also enhances the deraining prowess remarkably when real-world dataset is included.
Wu Ran, Peirong Ma, Zhiquan He, Hao Ren 0002, Hong Lu 0001
ICLR4
2024 Feature decoupling and reorganization network for single image deraining
Yunrui Cheng, Junjian Huang, Hao Ren 0002, Wu Ran, Hong Lu 0001
Multim. Syst.3
2024 Context-aware coarse-to-fine network for single image desnowing
Yunrui Cheng, Hao Ren 0002, Hong Lu 0001
Multim. Tools Appl.2
2024 Fully Unsupervised Domain-Agnostic Image Retrieval
abstract
Recent research in cross-domain image retrieval has focused on addressing two challenging issues: handling domain variations in the data and dealing with the lack of sufficient training labels. However, these problems have often been studied separately, limiting the practicality and significance of the research outcomes. The existing cross-domain setting is also restricted to cases where domain labels are known during training, and all samples have semantic category information or instance correspondences. In this paper, we propose a novel approach to address a more general and practical problem:fully unsupervised domain-agnostic image retrievalunder the domain-unknown setting, where no annotations are provided. Our approach tackles both thedomain variationandmissing labelschallenges simultaneously. We introduce a new fully unsupervised One-Shot Synthesis-based Contrastive learning method (termed OSSCo) to project images from different data distributions into a shared feature space for similarity measurement. To handle the domain-unknown setting, we propose One-Shot unpaired image-to-image Translation (OST) between a randomly selected one-shot image and the rest of the training images. By minimizing the global distance between the original images and the generated images from OST, the model learns domain-agnostic representations. To address the label-unknown setting, we employ contrastive learning with a synthesis-based transform module from the OST training. This allows for effective representation learning without any annotations or external constraints. We evaluate our proposed method on diverse datasets, and the results demonstrate its effectiveness. Notably, our approach achieves comparable performance to current state-of-the-art supervised methods.
Ziqiang Zheng, Hao Ren 0002, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.2
2024 Real-Time Attentive Dilated U-Net for Extremely Dark Image Enhancement
abstract
Images taken under low-light conditions suffer from poor visibility, color distortion, and graininess, all of which degrade the image quality and hamper the performance of downstream vision tasks, such as object detection and instance segmentation in the field of autonomous driving, making low-light enhancement an indispensable basic component of high-level visual tasks. Low-light enhancement aims to mitigate these issues, and has garnered extensive attention and research over several decades. The primary challenge in low-light image enhancement arises from the low signal-to-noise ratio caused by insufficient lighting. This challenge becomes even more pronounced in near-zero lux conditions, where noise overwhelms the available image information. Both traditional image signal processing pipeline and conventional low-light image enhancement methods struggle in such scenarios. Recently, deep neural networks have been used to address this challenge. These networks take unmodified RAW images as input and produce the enhanced sRGB images, forming a deep learning based image signal processing pipeline. However, most of these networks are computationally expensive and thus far from practical use. In this article, we propose a lightweight model called attentive dilated U-Net (ADU-Net) to tackle this issue. Our model incorporates several innovative designs, including an asymmetric U-shape architecture, dilated residual modules for feature extraction, and attentive fusion modules for feature fusion. The dilated residual modules provide strong representative capability, whereas the attentive fusion modules effectively leverage low-level texture information and high-level semantic information within the network. Both modules employ a lightweight design but offer significant performance gains. Extensive experiments demonstrate that our method is highly effective, achieving an excellent balance between image quality and computational complexity—that is, taking less than 4ms for a high-definition 4K image on a single GTX 1080Ti GPU and yet maintaining competitive visual quality. Furthermore, our method exhibits pleasing scalability and generalizability, highlighting its potential for widespread applicability.
Junjian Huang, Hao Ren 0002, Chuanlu Lv, Changyong Xie, Hong Lu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Instance-Aware Diffusion Implicit Process for Box-Based Instance Segmentation
abstract
The diffusion model has demonstrated impressive performance in image generation, but its potential for discriminative tasks such as instance segmentation remains unexplored. In this paper, we propose an Instance-aware Diffusion Implicit Process (IDIP) framework for instance segmentation based on boxes. During training, IDIP diffuses ground-truth boxes across various time steps, extracting corresponding Region of Interest (RoI) features. Dynamic convolution is then used to predict boxes and categories for each RoI, and the mask head generates masks from these predictions. During inference, IDIP iteratively refines randomly generated boxes with the denoising diffusion implicit model, while the mask head derives final masks from RoIs based on the refined boxes. Our method surpasses existing approaches on the COCO benchmark, requiring fewer training steps and less memory resources due to its dynamic design and instance-aware characteristic.
Hao Ren 0002, Xingsong Liu, Junjian Huang, Ru Wan, Jian Pu, Hong Lu 0001
ECAI1
2023 Weakly-supervised Temporal Action Localization with Adaptive Clustering and Refining Network
abstract
Weakly-supervised temporal action localization task aims to localize temporal boundaries of action instances by using only video-level labels. Existing methods primarily adopt Multi-Instance-Learning (MIL) scheme to handle this task. The effectiveness of MIL scheme depends heavily on the selection of top-k action snippets, which is unstable and requires manual tuning. To address these deficiencies, we propose an Adaptive Clustering and Refining Network (ACRNet). Specifically, we present an action-aware clustering strategy that is adaptable and requires no manual tuning to separate action and background snippets of diverse videos based on intra-class activation distribution. And a cluster refining step is included to eliminate false action snippets by considering inter-class activation distribution, which greatly improves robustness and localization accuracy. Extensive experiments on THUMOS14, ActivityNet 1.2&1.3 benchmarks show that our method achieves state-of-the-art performance.
Hao Ren 0002, Wu Ran, Xingson Liu, Hong Lu 0001, Cheng Jin 0001
ICME1
2023 Weakly-Supervised Temporal Action Localization with Regional Similarity Consistency
Hao Ren 0002, Hong Lu 0001, Cheng Jin 0001
MMM (1)2
2023 DaCo: domain-agnostic contrastive learning for visual place recognition
Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001
Appl. Intell.1
2023 ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image Retrieval
abstract
The huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset).
Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung
IEEE Trans. Circuits Syst. Video Technol.1
2022 Weakly-Supervised Temporal Action Localization with Multi-Head Cross-Modal Attention
Hao Ren 0002, Wu Ran, Hong Lu 0001, Cheng Jin 0001
PRICAI (3)1
2022 Energy-Guided Feature Fusion for Zero-Shot Sketch-Based Image Retrieval
Hao Ren 0002, Ziqiang Zheng, Hong Lu 0001
Neural Process. Lett.1
2022 Compositional coding capsule network with k-means routing for text classification
Hao Ren 0002, Hong Lu 0001
Pattern Recognit. Lett.1