Ben Wang 0005

dblp:53/5843-5 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-0880-9264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 40% Transfer learning and domain adaptation · 20% Efficient and distributed learning · 20%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
domain generalization
0.812024
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation › transfer learning for semantic segmentation
domain generalized semantic segmentation
0.812024
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation · CVPR 2024
Natural language and speech › Language models and text generation › large language model training › language model pretraining
masked pre-training
0.812024
Masked Pre-training Enables Universal Zero-shot Denoiser · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation · CVPR 2024
Image and video processing › image restoration
image denoising
0.812024
Masked Pre-training Enables Universal Zero-shot Denoiser · NeurIPS 2024
Image and video processing › image restoration › image denoising
zero-shot denoising
0.812024
Masked Pre-training Enables Universal Zero-shot Denoiser · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

masked pre-training · 1.5iterative filling · 1.5vision foundation model · 0.8trainable tokens · 0.8fine-tuning · 0.8
YearPublicationVenuePosition
2025 Seed Optimization With Frozen Generator for Superior Zero-Shot Low-Light Image Enhancement
abstract
In this work, we observe that the generators, which are pre-trained on massive natural images, inherently hold the promising potential for superior low-light image enhancement against varying scenarios. Specifically, for the low-light image enhancement process of a single image, we introduce the pre-trained generators to restore the details and colors degraded by low-light conditions, thereby improving the visual effect. Taking one step further, we introduce a novel optimization strategy, which backpropagates the gradients to the input seeds rather than the parameters of the low-light image enhancement model, thus intactly retaining the generative knowledge learned from natural images and achieving faster convergence speed. Benefiting from the pre-trained knowledge and seed-optimization strategy, the low-light image enhancement model can significantly regularize the visibility and fidelity of the enhanced result, thus rapidly generating high-quality images without training on any low-light dataset. Extensive experiments on various benchmarks demonstrate the effectiveness of the proposed method, showing its potential advantages over numerous state-of-the-art methods both qualitatively and quantitatively.
Yuxuan Gu 0001, Yi Jin 0002, Ben Wang 0005, Zhixiang Wei, Xiaoxiao Ma 0006, Haoxuan Wang 0004, Pengyang Ling, Huaian Chen, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.3
2025 Data and Prior-Driven Low-Light Enhancement Boosting the Visibility of Imaging Systems
Huaian Chen, Ben Wang 0005, Zhixiang Wei, Yi Jin 0002, Enhong Chen
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
abstract
In this paper, we first assess and harness various Vision Foundation Models (VFMs) in the context of Domain Generalized Semantic Segmentation (DGSS). Driven by the motivation that Leveraging Stronger pre-trained models and Fewer trainable parameters for Superior generalizability, we introduce a robust fine-tuning approach, namely “Rein”, to parameter-efficiently harness VFMs for DGSS. Built upon a set of trainable tokens, each linked to distinct instances, Rein precisely refines and forwards the feature maps from each layer to the next layer within the backbone. This process produces diverse refinements for different categories within a single image. With fewer trainable parameters, Rein efficiently fine-tunes VFMs for DGSS tasks, surprisingly surpassing full parameter fine-tuning. Extensive experiments across various settings demonstrate that Rein significantly outperforms state-of-the-art methods. Remarkably, with just an extra 1% of trainable parameters within the frozen backbone, Rein achieves a mIoU of 78.4% on the Cityscapes, without accessing any real urban-scene datasets. Code is available at https://github.com/w1oves/Rein.git.
Zhixiang Wei, Lin Chen 0026, Yi Jin 0002, Xiaoxiao Ma 0006, Pengyang Ling, Ben Wang 0005, Huaian Chen, Jinjin Zheng
CVPR7
2024 Masked Pre-training Enables Universal Zero-shot Denoiser
abstract
In this work, we observe that model trained on vast general images via masking strategy, has been naturally embedded with their distribution knowledge, thus spontaneously attains the underlying potential for strong image denoising. Based on this observation, we propose a novel zero-shot denoising paradigm, i.e., $\textbf{M}$asked $\textbf{P}$re-train then $\textbf{I}$terative fill ($\textbf{MPI}$). MPI first trains model via masking and then employs pre-trained weight for high-quality zero-shot image denoising on a single noisy image. Concretely, MPI comprises two key procedures: $\textbf{1) Masked Pre-training}$ involves training model to reconstruct massive natural images with random masking for generalizable representations, gathering the potential for valid zero-shot denoising on images with varying noise degradation and even in distinct image types. $\textbf{2) Iterative filling}$ exploits pre-trained knowledge for effective zero-shot denoising. It iteratively optimizes the image by leveraging pre-trained weights, focusing on alternate reconstruction of different image parts, and gradually assembles fully denoised image within limited number of iterations. Comprehensive experiments across various noisy scenarios underscore the notable advances of MPI over previous approaches with a marked reduction in inference time.
Xiaoxiao Ma 0006, Zhixiang Wei, Yi Jin 0002, Pengyang Ling, Ben Wang 0005, Junkang Dai, Huaian Chen
NeurIPS6
2024 Collaborative Filter Pruning for Efficient Automatic Surface Defect Detection
abstract
Surface defect detection is a critical task in industrial production, and numerous methods have been proposed to achieve high detection accuracy. Although deep-learning-based approaches have achieved state-of-the-art (SOTA) performances, their vast computational cost and high memory footprint prevent their deployment in resource-constrained environments. To address this problem, we propose a collaborative filter pruning method for the defect detection model, which significantly reduces the number of required calculations and parameters while maintaining high performance, even in cases with tasks suffering from the class imbalance problem. Our method aims to obtain lightweight pruned models by removing unimportant filters according to their importance evaluated by both structural similarity and detail richness of corresponding feature maps. Moreover, to improve the performance of pruned models, we propose a knowledge-fused fine-tuning approach that fuses the knowledge derived from two teacher networks to look after both representation learning and classifier learning, alleviating the class imbalance problem. Experimental results on four public datasets demonstrate that the proposed approach performs favorably relative to the SOTA methods. In particular, the proposed method achieves 39× and 59× parameter compression for VGG-16 and ResNet-50, respectively, on the NEU-CLS dataset, with a very small detection accuracy loss (<0.2%).
Haoxuan Wang 0004, Xin Fan 0005, Pengyang Ling, Ben Wang 0005, Huaian Chen, Yi Jin 0002
IEEE Trans. Ind. Informatics4
2023 BRAS: Bidirectional Reflectance Adjustment Strategy for 3-D Reconstruction of Mirror-Like Surface
abstract
A mirror-like surface (MLS) reflects highlight, aggravating the image saturation in structured light 3-D reconstruction systems and precluding defect detection based on reconstructed 3-D profiles. Previous studies have focused on a strategy of limiting the luminous flux entering the camera. However, the high intensity of the reflected highlight forces the limitation to be strengthened, which heavily reduces the modulation in the images for the 3-D reconstruction. Therefore, we propose a new strategy to adjust the source of the reflected highlight, i.e., the bidirectional reflectance (BR) of the MLS, which fundamentally suppresses the highlight and removes the limitation on the luminous flux. To execute the proposed strategy, an unfixed view structured light system (UVSLS) is established. The UVSLS converts the reflection viewer from the real camera to the virtual camera, realizing the flexible adjustment of the MLS BR. Finally, an MLS 3-D reconstruction framework is constructed to obtain the 3-D profile of the MLS. Experiments demonstrate that the proposed framework reduces the saturated pixels by 73.72% compared with the conventional method. Compared with the previous methods, the saturated pixels are reduced by an average of 47.62%.
Ben Wang 0005, Yabing Zheng, Minghui Duan, Xin Fan 0005, Yi Jin 0002, Jinjin Zheng
IEEE Trans. Ind. Informatics2