Shuocheng Wang

dblp:284/8256 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Area-latency-balanced Hardware Design for the Reference Pixel Management of VVC Intra Coding
Chengkang Huang, Leilei Huang, Taoyu Zhang, Shuocheng Wang, Yibo Fan
ISCAS4
2026 Controllable image-Guided generation via dynamic gaussian spectral modulation
Shuocheng Wang, Yuanbo Xing, Mengyuan Ge
Expert Syst. Appl.1
2026 A 77.9%-Cycle-Reduced Bubble-Removing Strategy for Hardware RDO Supporting QTMTT in VVC
abstract
The introduction of the Versatile Video Coding (VVC) standard is dedicated to meeting the increasing demand for high-resolution and high-quality video. However, the novel partitioning method named Quad-Tree Plus Multi-Type Tree (QTMTT) significantly increases computational complexity and data dependency, resulting in more pipeline bubbles, lower hardware efficiency, and degraded throughput. To address this issue, we analyze the hardware data dependencies, categorize four different types of pipeline bubbles, and propose a bubble-removing strategy for the Rate Distortion Optimization (RDO) module with MTT depth of 1. To be more specific, we first propose an efficient partition scheduling scheme based on the characteristics of QTMTT partitions. Then we redesign the transpose memory used in 2D transformation to efficiently handle the blocks of different sizes introduced by QTMTT. These two strategies achieve a 42.3% reduction in hardware cycles. In addition, for I frames, we further propose a hardware-oriented partition pruning algorithm that can co-operate with the proposed architecture, achieving a 42.3%~77.9% reduction in hardware cycles with only 0%~1.21% BD-Rate loss compared to the VTM-23.4. The proposed hardware architecture is implemented in GF 28nm technology, supporting up to 4K@40fps throughput at 500MHz with a hardware cost of only 3259 K gates and 63.47 KB on-chip memory, demonstrating competitive compression performance, high hardware efficiency, and outstanding throughput.
Chengkang Huang, Leilei Huang, Taoyu Zhang, Wei Li 0257, Shuocheng Wang, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.6
2026 One-Iteration ISP Controller for Real-Time Machine Vision
abstract
Conventional image signal processing (ISP) control algorithms based on human visual perception are insufficient for the demands of modern machine vision systems. Although recent learning-based methods have attempted to adapt ISP hyperparameters for specific vision tasks, their high latency and hardware cost of iterative optimization hinder deployment on edge devices such as those used in autonomous driving. To address these limitations, this paper proposes a real-time machine vision system through algorithm–hardware co-design. First, we introduce a one-iteration learning framework to avoid iterative optimization, significantly reducing latency for real-time use. Second, we propose a hardware-friendly controller, RasterNet, specifically tailored for raster-scanning sensor dataflow, eliminating redundant computation. Third, we present a pipelined ISP controller architecture incorporating branch and chroma time division multiplexing techniques to minimize the number of processing elements, achieving a compact and efficient design. Experiments demonstrate that the proposed system achieves superior object detection accuracy on resource-constrained platforms, with real-time performance reaching 70 FPS on FPGA and 224 FPS on ASIC implementations.
Zhijian Hao, Qi Zheng 0004, Ruoxi Zhu, Shuocheng Wang, Honglei Chen, Wenzhong Bao, Hongkai Xiong, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.5
2026 A Dual-Generalization Low-Light Enhancement Framework for Capsule Endoscopy Image Restoration and Segmentation
abstract
In recent years, deep learning technology has automated the diagnosis of gastrointestinal (GI) tract disease, enabling doctor-machine collaborative diagnosis. However, the images captured by wireless capsule endoscopy (WCE) easily suffer from varying brightness levels of low-light degradation due to the complex structure of GI tract and the limitations of the light source, which impacts both human and machine diagnostic accuracy. Moreover, images may contain varying degrees of structural and semantic details even under a similar brightness level, which still can compromise segmentation accuracy. To address these issues, we propose a dual-generalization framework for low-light WCE images. Our framework includes an Image Guidance and Laplacian Fusion Module (IGLFM), a Brightness Level Generalization Module (BLGM) and a Wavelet Segmentation Generalization Module (WSGM). IGLFM and BLGM can restore low-light images across different brightness levels and WSGM can enhance segmentation accuracy by generalizing the varying degrees of details across images. With BLGM and WSGM, our framework enables two aspects of generalization: generalization to input images with different brightness levels and generalization to images with varying detail levels. Extensive experiments demonstrate that our method achieves significant performance under varying brightness levels and improvements in segmentation accuracy, surpassing the existing state-of-the-art (SOTA) method with gains of 4.70 dB / 0.022 (PSNR/SSIM) on Kvasir-Capsule dataset and 1.61 dB / 0.018 on RLE dataset. WSGM consistently improves segmentation accuracy across six popular networks, achieving up to + 4.7% mIoU and + 5.3% Dice improvements on RLE dataset. Our code will be available at https://github.com/superwsc/Dual-Gen-Frame.
Shuocheng Wang, Ruoxi Zhu, Chengkang Huang, Minge Jing, Yibo Fan
IEEE Trans. Medical Imaging1
2025 MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware System
abstract
The rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems.
Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan
DAC6
2025 A Fast Saturation Based Dehazing Framework with Accelerated Convolution and Attention Block
abstract
Real-time image dehazmg is crucial for applications such as autonomous driving, surveillance, and remote sensing, where haze can significantly reduce visibility. However, many deep learning algorithms are hindered by large model sizes, making real-time processing difficult to achieve. Several fast and lightweight dehazing networks rely on estimating K(x), but they often fail to deliver satisfying performance. In this paper, we present a novel fast dehazing framework built upon the saturation-based algorithm. We design a new convolution module called Feature Extraction Partial Convolution (FEPC), which is faster and achieves better performance than the vanilla 3×3 convolution. Additionally, we fully leverage the information redundancy between feature map channels by dividing it into two parts along the channel dimension and designing a Self-Cross Attention Block (SCAB). The reduction in channel count significantly reduces computational load and improves the framework’s speed. Through extensive experiments, our method demonstrates not only a fast inference speed but also superior dehazing performance, providing a promising solution for real-time practical deployment. Our code will be available at https://github.coni/superwscZFSB-Dehazing-Framework.
Shuocheng Wang, Yilian Zhong, Ruoxi Zhu, Jiazheng Lian, Yibo Fan
ICASSP1
2025 Simplified one-sided Image-to-Image Translation with Reconstruction-Constrained Generative Adversarial Networks
abstract
The utilization of generative adversarial networks (GANs) for image-to-image translation has undergone extensive scholarly exploration. Nonetheless, prevailing unidirectional image translation tasks remain intricate and computationally inefficient due to the employment of complex computational methods and supplementary network modules, thereby imposing a significant computational burden. This study suggests integrating reconstruction constraints as a surrogate for complex ones, reducing computational burden. By skillfully combining these with adversarial losses, the network performs image translation more effectively. To enhance both translational and reconstruction competencies, we propose using the discriminator’s encoding mechanism to retain the image’s attributes. This approach results in a simplified yet powerful unidirectional image translation model, proven superior through comparative analysis with various GAN-based models. Additionally, the practical efficacy of our model is empirically verified through systematic experimentation.
Shuocheng Wang, Mengyuan Ge, Yingdong Wang
ICASSP1
2025 DLVQA: A Dynamic Loss Approach For Visual Question Answering with Language Biases
abstract
This study tackles the language-dependent bias in Visual Question Answering (VQA) models, where models tend to rely heavily on the semantic relationship between the question and predefined answers, often at the cost of a deeper understanding of the image content. While existing methods have focused on improving model architectures and applying data augmentation, few have explored solutions at the level of answer feature representation. To address this, we propose a novel dynamic loss function framework aimed at mitigating linguistic bias from the feature space perspective. The framework consists of two key components: (1) an answer-frequency-based margining mechanism, which balances the training weights of answers with different frequencies based on their statistical occurrence in large-scale datasets, and (2) an adaptive margining component that dynamically detects and corrects potential biases, considering the unique nature of question branches. This fusion of frequency-based and adaptive margins allows the model to better regulate the learning of high- and low-frequency answers. Extensive experiments show that our approach significantly enhances the VQA model’s performance, achieving an average improvement of 21% on two representative VQA-CP datasets, thereby demonstrating the effectiveness of addressing language bias through answer feature learning.
Shuocheng Wang
ICME1
2025 A contrastive self-supervised learning method for source-free EEG emotion recognition
Yingdong Wang, Qunsheng Ruan, Shuocheng Wang
User Model. User Adapt. Interact.4
2024 Improving High-Frequency Detail Handling in Conditional Image Generation via Iterative Denoising in Diffusion Models
abstract
Conditional image generation has long been a critical task in computer vision. In recent years, diffusion models have surpassed Generative Adversarial Networks (GANs) in unsupervised image generation, achieving state-of-the-art results. However, challenges remain in conditional image generation, particularly in effectively preserving high-frequency details. To address this issue, we introduce a novel approach that iteratively applies denoising operations within the diffusion process to better handle high-frequency details. Our method significantly enhances the image processing capabilities of Score-based Diffusion Models (SBDMs) without requiring additional training, making it particularly effective for tasks such as image translation and cartoonization. By preserving semantic coherence while refining high-frequency details, our technique not only improves the quality of translated images but also broadens the potential applications of SBDMs across various image processing tasks.
Shuocheng Wang, Mengyuan Ge, Yingdong Wang
BIBM1
2024 Fast Adaptive Loop Filter Algorithm Based on the Optimization of Class Merging
abstract
Adaptive loop filter (ALF) is one of the new tools adopted in the next generation video coding standard Versatile Video Coding (VVC). ALF leads to a performance improvement of 2%∼6% at the expense of high computational complexity and long processing time. Especially when encoding, ALF accounts 5%∼20% for total encoding runtime. To solve this problem, this paper proposes a fast algorithm for ALF by optimizing the process of class merging, which is an important step of rate-distortion optimization (RDO) in ALF. Experimental results indicate that compared to the VVC Test Model (VTM-22.0), this method can reduce the ALF encoding runtime on average 48.16%, 60.65%, 56.15% under all-intra, low-delay P and random access configurations, with only 0.15%, 0.31%, 0.19% increase in luma Bjontegaard Delta-Bit Rate (BD-BR).
Chengkang Huang, Leilei Huang, Shuocheng Wang, Yibo Fan
VCIP4
2024 MI-EEG: Generalized model based on mutual information for EEG emotion recognition without adversarial training
Yingdong Wang, Shuocheng Wang, Xiqiao Fang, Qunsheng Ruan
Expert Syst. Appl.3
2023 LEAP: TrustZone Based Developer-Friendly TEE for Intelligent Mobile Apps
abstract
ARM TrustZone is widely deployed on commercial-off-the-shelf mobile devices for secure execution. However, many Apps cannot enjoy this feature because it brings many constraints to App developers. Previous works have been proposed to build a secure execution environment for developers on top of TrustZone. Unfortunately, these works are still not a fully-fledged solution for mobile Apps, especially for the emerging intelligent Apps. To this end, we propose LEAP, which is a lightweight developer-friendly TEE solution for mobile Apps. LEAP enables isolated codes to execute in parallel and access peripheral (e.g., mobile GPUs) with ease, flexibly manages system resources upon different workloads, and offers the auto DevOps tool to help developers prepare the codes running on it. We implement the LEAP prototype on the off-the-shelf ARM platform and conduct extensive experiments on it. The experimental results show that Apps can be adapted to run with LEAP easily and efficiently. Compared to the state-of-the-art work along this research line, LEAP can achieve an average 3.57× speedup in supporting intelligent Apps using mobile GPU acceleration.
Lizhi Sun, Shuocheng Wang, Hao Wu 0067, Yuhang Gong, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002
IEEE Trans. Mob. Comput.2
2021 Cross-subject EEG emotion classification based on few-label adversarial domain adaption
Yingdong Wang, Qunsheng Ruan, Shuocheng Wang, Chen Wang 0082
Expert Syst. Appl.4