Ruoxi Zhu

dblp:265/5033 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 One-Iteration ISP Controller for Real-Time Machine Vision
abstract
Conventional image signal processing (ISP) control algorithms based on human visual perception are insufficient for the demands of modern machine vision systems. Although recent learning-based methods have attempted to adapt ISP hyperparameters for specific vision tasks, their high latency and hardware cost of iterative optimization hinder deployment on edge devices such as those used in autonomous driving. To address these limitations, this paper proposes a real-time machine vision system through algorithm–hardware co-design. First, we introduce a one-iteration learning framework to avoid iterative optimization, significantly reducing latency for real-time use. Second, we propose a hardware-friendly controller, RasterNet, specifically tailored for raster-scanning sensor dataflow, eliminating redundant computation. Third, we present a pipelined ISP controller architecture incorporating branch and chroma time division multiplexing techniques to minimize the number of processing elements, achieving a compact and efficient design. Experiments demonstrate that the proposed system achieves superior object detection accuracy on resource-constrained platforms, with real-time performance reaching 70 FPS on FPGA and 224 FPS on ASIC implementations.
Zhijian Hao, Qi Zheng 0004, Ruoxi Zhu, Shuocheng Wang, Honglei Chen, Wenzhong Bao, Hongkai Xiong, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.4
2026 A Flexible Zero-Shot Approach to Tone Mapping via Structure-Preserving Diffusion Models
abstract
With the prevalence of high dynamic range (HDR) imaging, tone mapping techniques, which convert HDR images to high-quality standard dynamic range (SDR) images for display, have become increasingly important. However, obtaining paired HDR and high-quality SDR images is almost impossible, posing challenges to learning-based tone mapping methods. To address this issue, we propose a zero-shot tone mapping framework without requiring any HDR training samples. Our approach decomposes images into two components: structural information and tonal information. A diffusion-based mapping model taking the structural information as input is first trained in the high-quality SDR domain, then transferred to the HDR domain that has less readily available training data for inference, leveraging the equivalent distribution of the structural information across both domains. To preserve the original image’s structure, we modify the reverse sampling process and explicitly incorporate the original structural information into the intermediate results. To improve the image details, we introduce a dual-control network, enabling different conditional inputs to control different scales of the output. Additionally, we devise a flexible tone adjustment strategy, with a bunch of novel loss functions to modify the trained score function dynamically during reverse sampling, allowing users to customize the style of the generated image according to their preference during testing. Initially designed for tone mapping, our model can be applied to various tasks including image fusion, exposure correction, dehazing, etc., without retraining. Experimental results demonstrate that our approach surpasses previous state-of-the-art methods, indicating that it can serve as an effective, flexible and versatile solution to various tone-mapping tasks. Source code is available at https://github.com/ZSDM-HDR/Zero-Shot-Diffusion-HDR.
Ruoxi Zhu, Shusong Xu, Peiye Liu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Ming-e Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.1
2026 A Dual-Generalization Low-Light Enhancement Framework for Capsule Endoscopy Image Restoration and Segmentation
abstract
In recent years, deep learning technology has automated the diagnosis of gastrointestinal (GI) tract disease, enabling doctor-machine collaborative diagnosis. However, the images captured by wireless capsule endoscopy (WCE) easily suffer from varying brightness levels of low-light degradation due to the complex structure of GI tract and the limitations of the light source, which impacts both human and machine diagnostic accuracy. Moreover, images may contain varying degrees of structural and semantic details even under a similar brightness level, which still can compromise segmentation accuracy. To address these issues, we propose a dual-generalization framework for low-light WCE images. Our framework includes an Image Guidance and Laplacian Fusion Module (IGLFM), a Brightness Level Generalization Module (BLGM) and a Wavelet Segmentation Generalization Module (WSGM). IGLFM and BLGM can restore low-light images across different brightness levels and WSGM can enhance segmentation accuracy by generalizing the varying degrees of details across images. With BLGM and WSGM, our framework enables two aspects of generalization: generalization to input images with different brightness levels and generalization to images with varying detail levels. Extensive experiments demonstrate that our method achieves significant performance under varying brightness levels and improvements in segmentation accuracy, surpassing the existing state-of-the-art (SOTA) method with gains of 4.70 dB / 0.022 (PSNR/SSIM) on Kvasir-Capsule dataset and 1.61 dB / 0.018 on RLE dataset. WSGM consistently improves segmentation accuracy across six popular networks, achieving up to + 4.7% mIoU and + 5.3% Dice improvements on RLE dataset. Our code will be available at https://github.com/superwsc/Dual-Gen-Frame.
Shuocheng Wang, Ruoxi Zhu, Chengkang Huang, Minge Jing, Yibo Fan
IEEE Trans. Medical Imaging3
2025 MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware System
abstract
The rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems.
Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan
DAC4
2025 A Fast Saturation Based Dehazing Framework with Accelerated Convolution and Attention Block
abstract
Real-time image dehazmg is crucial for applications such as autonomous driving, surveillance, and remote sensing, where haze can significantly reduce visibility. However, many deep learning algorithms are hindered by large model sizes, making real-time processing difficult to achieve. Several fast and lightweight dehazing networks rely on estimating K(x), but they often fail to deliver satisfying performance. In this paper, we present a novel fast dehazing framework built upon the saturation-based algorithm. We design a new convolution module called Feature Extraction Partial Convolution (FEPC), which is faster and achieves better performance than the vanilla 3×3 convolution. Additionally, we fully leverage the information redundancy between feature map channels by dividing it into two parts along the channel dimension and designing a Self-Cross Attention Block (SCAB). The reduction in channel count significantly reduces computational load and improves the framework’s speed. Through extensive experiments, our method demonstrates not only a fast inference speed but also superior dehazing performance, providing a promising solution for real-time practical deployment. Our code will be available at https://github.coni/superwscZFSB-Dehazing-Framework.
Shuocheng Wang, Yilian Zhong, Ruoxi Zhu, Jiazheng Lian, Yibo Fan
ICASSP4
2024 Zero-Shot Structure-Preserving Diffusion Model for High Dynamic Range Tone Mapping
abstract
Tone mapping techniques, aiming to convert high dynamic range (HDR) images to high-quality low dynamic range (LDR) images for display, play a more crucial role in real-world vision systems with the increasing application of HDR images. However, obtaining paired HDR and high-quality LDR images is difficult, posing a challenge to deep learning based tone mapping methods. To over-come this challenge, we propose a novel zero-shot tone mapping framework that utilizes shared structure knowl-edge, allowing us to transfer a pre-trained mapping model from the LDR domain to HDR fields without paired training data. Our approach involves decomposing both the LDR and HDR images into two components: structural in-formation and tonal information. To preserve the original image's structure, we modify the reverse sampling process of a diffusion model and explicitly incorporate the struc-ture information into the intermediate results. Additionally, for improved image details, we introduce a dual-control network architecture that enables different types of conditional inputs to control different scales of the output. Experimental results demonstrate the effectiveness of our approach, surpassing previous state-of-the-art methods both qualitatively and quantitatively. Moreover, our model ex-hibits versatility and can be applied to other low-level vi-sion tasks without retraining. The code is available at https://github.com/ZSDM-HDRIZero-Shot-Diffusion-HDR.
Ruoxi Zhu, Shusong Xu, Peiye Liu, Sicheng Li 0001, Yanheng Lu, Dimin Niu, Zihao Liu 0015, Zihao Meng, Zhiyong Li 0016, Xinhua Chen, Yibo Fan
CVPR1
2024 Auto-ISP: An Efficient Real-Time Automatic Hyperparameter Optimization Framework for ISP Hardware System
abstract
Image Signal Processor (ISP) is widely used in intelligent edge devices across various scenarios. The intricate and time-consuming tuning process demands substantial expertise. Current AI-based auto-tuning operates discretely offline, relying on predefined scenes with human intervention, leading to inconvenient manipulation, with potentially fatal impacts on downstream tasks in unforeseen scenes. We propose a real-time automatic hyperparameter optimization ISP hardware system to address real-world scenarios. Our design features a tri-step framework and a hardware accelerator, demonstrating superior performance in human and computer vision tasks, even in real-time unforeseen scenes. Experiments showcase its practicality, achieving 1080P@75FPS/240FPS in FPGA/ASIC, respectively.
Zihao Liu 0015, Ruoxi Zhu, Qi Zheng 0004, Zhijian Hao, Tao Liu 0023, Jun Tao 0001, Yibo Fan
DAC4
2024 Multi-Weather Degradation-Aware Transformer for Image Restoration
abstract
Restoring images under different adverse weather conditions with a single model is practical in many applications. Most existing weather restoration approaches are only able to handle a specific type of degradation, which is often insufficient in real-world scenarios where the weather type is unknown. In this paper, we propose a holistic solution to solve multiple weather degradations using a single model. Specifically, we build a weather-type aware Transformer, an efficient architecture that can restore images degraded by different adverse weathers with the same set of parameters. For model training, we first use contrastive loss to train an auxiliary hypernetwork capable of extracting content-independent, distortion-aware feature embeddings. Guided by these weather-dependent features, the image restoration Transformer can adaptively modulate its parameters using hypernetworks and feature-wise linear modulation blocks, conducting both local and global operations adaptively for images with different degradations. Qualitative and quantitative results on the multi-weather benchmark demonstrate that our model achieves significant improvements compared with previous state-of-the-arts, with even less computational cost.
Ruoxi Zhu, Minfeng Wu, Xiankui Xiong, Xuanpeng Zhu, Yibo Fan
ICASSP1
2024 MWFormer: Multi-Weather Image Restoration Using Degradation-Aware Transformers
abstract
Restoring images captured under adverse weather conditions is a fundamental task for many computer vision applications. However, most existing weather restoration approaches are only capable of handling a specific type of degradation, which is often insufficient in real-world scenarios, such as rainy-snowy or rainy-hazy weather. Towards being able to address these situations, we propose a multi-weather Transformer, or MWFormer for short, which is a holistic vision Transformer that aims to solve multiple weather-induced degradations using a single, unified architecture. MWFormer uses hyper-networks and feature-wise linear modulation blocks to restore images degraded by various weather types using the same set of learned parameters. We first employ contrastive learning to train an auxiliary network that extracts content-independent, distortion-aware feature embeddings that efficiently represent predicted weather types, of which more than one may occur. Guided by these weather-informed predictions, the image restoration Transformer adaptively modulates its parameters to conduct both local and global feature processing, in response to multiple possible weather. Moreover, MWFormer allows for a novel way of tuning, during application, to either a single type of weather restoration or to hybrid weather restoration without any retraining, offering greater controllability than existing methods. Our experimental results on multi-weather restoration benchmarks show that MWFormer achieves significant performance improvements compared to existing state-of-the-art methods, without requiring much computational cost. Moreover, we demonstrate that our methodology of using hyper-networks can be integrated into various network architectures to further boost their performance. The code is available at: https://github.com/taco-group/MWFormer.
Ruoxi Zhu, Zhengzhong Tu, Alan C. Bovik, Yibo Fan
IEEE Trans. Image Process.1
2023 Luminance-Preserving Visible and Near-Infrared Image Fusion Network with Edge Guidance
abstract
Near-infrared (NIR) images and visible (VIS) images can provide mutually complementary information for each other, thus the fusion of the two modalities can create images of high quality even in adverse conditions. However, the luminance of NIR and VIS images may be inconsistent in some regions, resulting in color distortion and unrealistic appearance in the fused images. The existing methods perform poorly at luminance retention. Aiming at the problem and based on deep learning framework, we propose an edge-guided method which can be applied to the image fusion network. Edge maps are utilized as prior knowledge of images to boost the performance of the neural network. Additionally, we propose a luminance-preserving loss function combined with max-edge loss to further improve the image quality. Experimental results show the superiority of our method.
Ruoxi Zhu, Yi Ling, Xiankui Xiong, Dong Xu 0015, Xuanpeng Zhu, Yibo Fan
ICIP1
2023 Machine Learning-based Intrusion Detection for Smart Grid Computing: A Survey
abstract
Machine learning (ML)-based intrusion detection system (IDS) approaches have been significantly applied and advanced the state-of-the-art system security and defense mechanisms. In smart grid computing environments, security threats have been significantly increased as shared networks are commonly used, along with the associated vulnerabilities. However, compared to other network environments, ML-based IDS research in a smart grid is relatively unexplored, although the smart grid environment is facing serious security threats due to its unique environmental vulnerabilities. In this article, we conducted an extensive survey on ML-based IDS in smart grids based on the following key aspects: (1) The applications of the ML-based IDS in transmission and distribution side power components of a smart power grid by addressing its security vulnerabilities; (2) dataset generation process and its usage in applying ML-based IDSs in the smart grid; (3) a wide range of ML-based IDSs used by the surveyed papers in the smart grid environment; (4) metrics, complexity analysis, and evaluation testbeds of the IDSs applied in the smart grid; and (5) lessons learned, insights, and future research directions.
Nitasha Sahani, Ruoxi Zhu, Jin-Hee Cho, Chen-Ching Liu
ACM Trans. Cyber Phys. Syst.2