Zhuoran Zheng

dblp:293/8326 · DBLP profile ↗
← Back
48ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0002-3617-6513ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 25 since 2021Artificial intelligence and machine learning · 24 · 5 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DCA-LUT: Deep Chromatic Alignment with 5D LUT for Purple Fringing Removal
abstract
Purple fringing, a persistent artifact caused by Longitudinal Chromatic Aberration (LCA) in camera lenses, has long degraded the clarity and realism of digital imaging. Traditional solutions rely on complex and expensive apochromatic (APO) lens hardware and the extraction of handcrafted features, ignoring the data-driven approach. To fill this gap, we introduce DCA-LUT, the first deep learning framework for purple fringing removal. Inspired by the physical root of the problem-the spatial misalignment of RGB color channels due to lens dispersion, we introduce a novel Chromatic-Aware Coordinate Transformation (CA-CT) module, learning an image-adaptive color space to decouple and isolate fringing into a dedicated dimension. This targeted separation allows the network to learn a precise "purple fringe channel," which then guides the accurate restoration of the luminance channel. The final color correction is performed by a learned 5D Look-Up Table (5D LUT), enabling efficient and powerful non-linear color mapping. To enable robust training and fair evaluation, we constructed a large-scale synthetic purple fringing dataset (PF-Synth). Extensive experiments in synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance in purple fringing removal.
Jialang Lu, Shuning Sun, Pu Wang 0008, Chen Wu 0006, Feng Gao 0005, Lina Gong, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
AAAI9
2026 CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal
abstract
Purple flare, a diffuse chromatic aberration artifact commonly found around highlight areas, severely degrades the tone transition and color of the image. Existing traditional methods are based on hand-crafted features, which lack flexibility and rely entirely on fixed priors, while the scarcity of paired training data critically hampers deep learning. To address this issue, we propose a novel network built upon decoupled HSV Look-Up Tables (LUTs). The method aims to simplify color correction by adjusting the Hue (H), Saturation (S), and Value (V) components independently. This approach resolves the inherent color coupling problems in traditional methods. Our model adopts a two-stage architecture: First, a Chroma-Aware Spectral Tokenizer (CAST) converts the input image from RGB space to HSV space and independently encodes the Hue (H) and Value (V) channels into a set of semantic tokens describing the Purple flare status; second, the HSV-LUT module takes these tokens as input and dynamically generates independent correction curves (1D-LUTs) for the three channels H, S, and V. To effectively train and validate our model, we built the first large-scale purple flare dataset with diverse scenes. We also proposed new metrics and a loss function specifically designed for this task. Extensive experiments demonstrate that our model not only significantly outperforms existing methods in visual effects but also achieves state-of-the-art performance on all quantitative metrics.
Pu Wang 0008, Shuning Sun, Jialang Lu, Chen Wu 0006, Youshan Zhang, Chenggang Shan, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
AAAI10
2026 Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
Chen Wu 0006, Zhuoran Zheng, Jingyuan Xia, Weidong Jiang
ISCAS3
2026 Fusion requires interaction: a hybrid Mamba-transformer architecture for deep interactive fusion of multi-modal images
Wenxiao Xu, Chen Wu 0006, Qiyuan Yin, Zhuoran Zheng, Daqing Huang
Expert Syst. Appl.5
2026 Defense against unauthorized distillation in image restoration via feature space perturbation
Zhuoran Zheng, Chen Lyu 0001
Neurocomputing2
2026 UHD image dehazing via anDehazeFormer with atmospheric-aware KV cache
Pu Wang 0008, Zhixuan Mao, Wenhao Li 0006, Liubing Hu, Dianjie Lu, Guijuan Zhang, Youshan Zhang, Zhuoran Zheng
Neurocomputing8
2026 EPIDIR: Integrating external frequency prompts and internal statistical decoupling for infrared image restoration
Wenxiao Xu, Qiyuan Yin, Zhuoran Zheng, Daqing Huang
Neurocomputing3
2026 Label distribution learning via implicit distribution representation
Zhuoran Zheng, Xin Su 0009, Chen Lyu 0001
Neurocomputing1
2026 CLIP2LE: A Label Enhancement Fair Representation Method via CLIP
Pu Wang 0008, YinSong Xiong, Zhuoran Zheng
IEEE Trans. Big Data3
2026 Ultra-High-Definition Image Restoration via High-Frequency Enhanced Transformer
abstract
Transformer-based architectures exhibit substantial promise in the realm of ultra-high-definition (UHD) image restoration (IR). Nevertheless, they encounter significant challenges in maintaining high-frequency (HF) details, which are crucial for the reconstruction of texture. Conventional methods tackle computational complexity by significantly reducing the resolution (by a factor of 4 to 8). Moreover, the majority of high-frequency components are eliminated due to the inherent characteristics of self-attention mechanisms, as these mechanisms tend to naturally suppress high-frequency elements during non-local feature integration. This paper proposes a dual-branch transformer architecture that synergistically combines native-resolution HF preservation with efficient contextual modeling, named HiFormer. The high-resolution branch utilizes a directionally-sensitive large-kernel decomposition to effectively address anisotropic degradations with fewer parameters and applies depthwise separable convolutions for localized high-frequency (HF) information extraction. Concurrently, the low-resolution branch assimilates these localized HF elements using adaptive channel modulation to offset spectral losses induced by the inherent smoothing effect of self-attention. Comprehensive experiments across numerous UHD image restoration tasks reveal that our approach surpasses current leading methods in both quantitative metrics and qualitative analysis. The code is available at https://github.com/5chen/HiFormer.
Chen Wu 0006, Zhuoran Zheng, Weidong Jiang, Yuning Cui 0001, Jingyuan Xia
IEEE Trans. Circuits Syst. Video Technol.3
2026 TSFormer: Efficient Ultra-High-Definition Image Restoration via Trusted Min-p
abstract
Ultra-high-definition (UHD) image restoration is vital for applications demanding exceptional visual fidelity, yet existing methods often face a trade-off between restoration quality and efficiency, limiting their practical deployment. In this paper, we propose TSFormer, an all-in-one framework that integrates Trusted learning with Sparsification to boost both generalization capability and computational efficiency in UHD image restoration. The key to sparsification is that only a small amount of token movement is allowed within the model. To efficiently filter tokens, we use Min- $p$ with random matrix theory to quantify the uncertainty of tokens (lower trustworthiness), thereby improving the robustness of the model. Our model can run a 4K ( $3840\times 2160$ ) image in real time (40fps) with 3.38 M parameters. Extensive experiments demonstrate that TSFormer achieves state-of-the-art restoration quality while enhancing generalization and reducing computational demands. In addition, our token filtering method can be applied to other image restoration models to effectively accelerate inference and maintain performance.
Zhuoran Zheng, Pu Wang 0008, Liubing Hu, Xin Su 0009
IEEE Trans. Image Process.1
2026 Towards Ultra-High-Definition Image Deraining: A Benchmark and an Efficient Method
abstract
Despite significant advancements in image deraining, most existing methods are carried out on low-resolution images, leaving their effectiveness on high-resolution images uncertain. This limitation becomes even more pronounced with the rise of ultra-high-definition (UHD) imaging. In this paper, we tackle the challenge of UHD image deraining and introduce 4K-Rain13 k, the first large-scale UHD image deraining dataset, featuring 13,000 paired images at 4 K resolution. Leveraging this dataset, we conduct a benchmark study on existing methods for processing UHD images. To better address this task, we propose UDR-Mixer, an efficient and effective architecture tailored for UHD image deraining. Our model comprises two key components: a spatial feature rearrangement layer, which captures long-range dependencies in UHD images, and a frequency feature modulation layer, which enhances high-fidelity image reconstruction. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while maintaining lower model complexity. The source code and proposed dataset are available athttps://github.com/cschenxiang/UDR-Mixer.
Hongming Chen 0004, Xiang Chen 0015, Chen Wu 0006, Zhuoran Zheng, Jinshan Pan, Xianping Fu
IEEE Trans. Multim.4
2025 Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel Learning
abstract
With the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating high-quality UHD images on a resource-constrained device. To alleviate the limitations of extremely long-sequence inputs, inspired by the Large Language Model (LLM) for processing infinitely long texts, we propose a novel learning paradigm to achieve UHD multi-exposure dynamic scene image fusion on a single consumer-grade GPU, named Infinite Pixel Learning (IPL). The design of our approach comes from three key components: The first step is to slice the input sequences to relieve the pressure generated by the model processing the data stream; Second, we develop an attention cache technique, which is similar to the KV cache for infinite data stream processing; Finally, we design a method for attention cache compression to alleviate the storage burden of the cache on the device. In addition, we provide a new UHD benchmark to evaluate the effectiveness of our method. Extensive experimental results show that our method maintains high-quality visual performance while fusing UHD dynamic multi-exposure images in real-time (>40fps) on a single consumer-grade GPU.
Xingchi Chen, Zhuoran Zheng, Xuerui Li, Yuying Chen, Wenqi Ren
AAAI2
2025 LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt Tuning
abstract
Visual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tremendous domain gap leads to significant performance degradation. To address this issue, in this paper, we propose a novel robust UAV tracker termed LVPTrack, which conducts high quality label-aligned visual prompt tuning to adapt to various challenging weather conditions. Specifically, we first synthesize the sequential foggy and nighttime video frames to assist the model training. A domain adaptive teacher-student network is utilized to distill the hierarchical visual semantic of the target objects in cross-domain scenarios. Then we propose a target-aware pseudo-label voting (PLV) strategy to alleviate the target-level misalignment in the dual domains. Furthermore, we propose a dynamic aggregated prompt (DAP) module to facilitate the appearance variation adaptation of the target object in challenging scenarios. Extensive experiments demonstrate that our tracker achieves superior performance over existing state-of-the-art UAV trackers.
Hongjing Wu, Siyuan Yao, Feng Huang 0007, Linchao Zhang, Zhuoran Zheng, Wenqi Ren
AAAI6
2025 An Effective Wavelet Neural Network for Ultra-High-Definition Image Deraining
Zhuoran Zheng, Shao-Ping Lu, Yulu Yang
CGI (3)2
2025 ECSNN: Spiking Neural Networks for Efficient Exposure Correction in Endoscopy Imaging
abstract
The quality of endoscopic images is critical to the success of polyp segmentation, highlighting the need for accurate exposure correction in endoscopy. While traditional deep learning methods are effective, they demand substantial computational resources during inference. To address this, we propose the Endoscopic Exposure Correction Spiking Neural Network (ECSNN), an efficient framework designed for resource-limited devices. Our approach features a Positive Incentive Learning Module that reduces noise in input images. These enhanced features are then processed by U-Shape Networks (USNet), which leverages spiking neural networks to learn deep representations for exposure correction. Additionally, we introduce a Brightness Prompt Module consisting of two components: the Brightness Spike Encoding Module (BSEM), which encodes brightness information into spike signals, and the Brightness-Aware Prompt Block (BAPB), which adjusts exposure by guiding the network through brightness-aware attention. We evaluate ECSNN on the Endo4IE and ECSEG datasets, where it outperforms six state-of-the-art methods and demonstrates its practical utility in clinical diagnosis.
Jun Zhang 0011, Zhuoran Zheng, Jingang Zhang, Wenqi Ren
ICASSP2
2025 UniFlowRestore: A General Video Restoration Framework via Flow Matching and Prompt Guidance
Shuning Sun, Yu Zhang 0296, Chen Wu 0006, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
ACM Multimedia7
2025 ChebSpec-Net: Linear Spectral Graph Restoration for UHD Images
Xin Su 0009, Zhuoran Zheng, Jianshu Chao, Dapeng Ye, Zhicong Luo
PRCV (8)2
2025 Dropout the High-Rate Downsampling: A Novel Design Paradigm for UHD Image Restoration
abstract
With the popularization of high-end mobile devices, Ultra-high-definition (UHD) images have become ubiquitous in our lives. The restoration of UHD images is a highly challenging problem due to the exaggerated pixel count, which often leads to memory overflow during processing. Existing methods either downsample UHD images at a high rate before processing or split them into multiple patches for separate processing. However, high-rate downsampling leads to significant information loss, while patch-based approaches inevitably introduce boundary artifacts. In this paper, we propose a novel design paradigm to solve the UHD image restoration problem, called D2Net. D2Net enables direct full-resolution inference on UHD images without the need for high-rate downsampling or dividing the images into several patches. Specifically, we ingeniously utilize the characteristics of the frequency domain to establish long-range dependencies of features. Taking into account the richer local patterns in UHD images, we also design a multi-scale convolutional group to capture local features. Additionally, during the decoding stage, we dynamically incorporate features from the encoding stage to reduce the flow of irrelevant information. Extensive experiments on three UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring, show that our model achieves better quantitative and qualitative results than state-of-the-art methods.
Chen Wu 0006, Long Peng 0003, Dianjie Lu, Zhuoran Zheng
WACV5
2025 Complex mixer for MedMNIST classification decathlon
Shuning Sun, Xiuyi Jia, Zhuoran Zheng
Appl. Intell.3
2025 4K-HAZE: A dehazing benchmark with 4K resolution hazy and haze-free images
Xin Su 0009, Pengwen Dai, Zhuoran Zheng
Neurocomputing4
2025 Multi-Exposure image Fusion via distilled 3D LUT grid with editable mode
Xin Su 0009, Zhuoran Zheng, Jialing Yang
Neurocomputing2
2025 MixNet: Efficient global modeling for ultra-high-definition image restoration
Chen Wu 0006, Shuning Sun, Yu Zhang 0296, Zhuoran Zheng
Neurocomputing4
2025 Re-examine all-in-one image restoration: A catastrophic forgetting perspective
Chen Wu 0006, Pu Wang 0008, Zhuoran Zheng
Pattern Recognit. Lett.3
2025 Instance-Wise Privacy Preservation for All-in-One Image Restoration
Pu Wang 0008, Xin Su 0009, Zhuoran Zheng
IEEE Signal Process. Lett.3
2025 AgentPolyp: Accurate Polyp Segmentation via Image Enhancement Agent
abstract
Captured polyp images often suffer from degradation, such as dim lighting, blur, and overexposure. Direct segmentation is prone to artifact diffusion, which significantly degrades the performance of downstream segmentation algorithms and leads to inaccurate boundary delineation. Addressing these varied degradations requires a dynamic, intelligent process that diagnoses and applies targeted corrections. We present AgentPolyp, a novel framework driven by an intelligent agent that integrates CLIP-based semantic guidance and dynamic image enhancement with a lightweight segmentation network. The agent adaptively selects reinforcement learning strategies to perform context-aware denoising, contrast adjustment, and artifact reduction. This selection process is continuously optimized through a feedback loop that includes quality assessment, ensuring the optimization and enhancement of downstream segmentation. This approach addresses degradation complexity and feature compatibility issues, offering a deployable solution for endoscopic polyp analysis.
Pu Wang 0008, Guangwei Gao, Youshan Zhang, Zhuoran Zheng
IEEE Signal Process. Lett.5
2025 Adaptive Feature Selection Modulation Network for Efficient Image Super-Resolution
abstract
In the realm of image super-resolution, learning-based methods have made significant progress. However, limited computational resources still restrict their application. This prompts us to develop an efficient method for achieving effective image super-resolution. In this letter, we propose a novel adaptive feature selection modulation network (AFSMNet) tailored for efficient image super-resolution. Specifically, we design feature modulation blocks, which include the adaptive feature selection modulation (AFSM) module and the self-gating feed-forward network (SFN). The AFSM module dynamically computes the importance of each feature channel. For channels with differing levels of importance, we employ distinct processing strategies, thereby concentrating the computational resources of the network on the more critical features as much as possible. This approach facilitates the maintenance of a low computational cost without compromising performance. The SFN restricts the flow of irrelevant feature information within the network through a simple gating mechanism. In this way, our method achieves efficient and effective image super-resolution. Extensive experiment results show that the proposed method achieves a better trade-off between reconstruction performance and computational efficiency compared to the current state-of-the-art lightweight super-resolution methods.
Chen Wu 0006, Xin Su 0009, Zhuoran Zheng
IEEE Signal Process. Lett.4
2025 NSDSAM: Noise-Suppression-Driven SAM for Infrared Small Target Detection
abstract
Although Segment Anything Model (SAM) have recently achieved remarkable progress, their generalization capability in infrared small target detection remains limited due to the inherently high noise levels in infrared imagery. To preserve the generalization and noise suppression ability of the model, we propose an method called NSDSAM, a noise-suppression-driven approach that enhances SAM at both internal and external levels. Internally, we develop a Hybrid Adapter for suppressing the noise of feature maps, consisting of an MLP adapter and a self-attention adapter. The self-attention adapter first performs entropy-aware reconstruction of features from noisy inputs and employs a gating mechanism for soft-attention fusion, mitigating SAM’s sensitivity to noise. Externally, we design a Spatial-Frequency hybrid Module (SFHM) that jointly processes spatial and frequency domains to overcome the self-attention model’s bias toward low-frequency components, further strengthening the suppression of background clutter and noise. Extensive experiments on multiple infrared datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance in infrared small target detection. The project code is available upon acceptance.
Wenxiao Xu, Qiyuan Yin, Chen Wu 0006, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
IEEE Trans. Geosci. Remote. Sens.6
2025 M2Restore: Mixture-of-Experts-Based Mamba-CNN Fusion Framework for All-in-One Image Restoration
abstract
Natural images are often degraded by complex, composite degradations such as rain, snow, and haze, which adversely impact downstream vision applications. While existing image restoration efforts have achieved notable success, they are still hindered by two critical challenges: limited generalization across dynamically varying degradation scenarios and a suboptimal balance between preserving local details and modeling global dependencies. To overcome these challenges, we propose M2Restore, a novel Mixture-of-Experts (MoE)-based Mamba-CNN fusion framework for efficient and robust all-in-one image restoration. M2Restore introduces three key contributions: First, to boost the model's generalization across diverse degradation conditions, we exploit a CLIP-guided MoE gating mechanism that fuses task-conditioned prompts with CLIP-derived semantic priors. This mechanism is further refined via cross-modal feature calibration, which enables precise expert selection for various degradation types. Second, to jointly capture global contextual dependencies and fine-grained local details, we design a dual-stream architecture that integrates the localized representational strength of CNNs with the long-range modeling efficiency of Mamba. This integration enables collaborative optimization of global semantic relationships and local structural fidelity, preserving global coherence while enhancing detail restoration. Third, we introduce an edge-aware dynamic gating mechanism that adaptively balances global modeling and local enhancement by reallocating computational attention to degradation-sensitive regions. This targeted focus leads to more efficient and precise restoration. Extensive experiments across multiple image restoration benchmarks validate the superiority of M2Restore in both visual quality and quantitative performance. Code is available at https://github.com/yz-wang/M2Restore.
Yongzhen Wang 0001, Zhuoran Zheng, Xiao-Ping Zhang 0002, Mingqiang Wei
IEEE Trans. Image Process.3
2025 Group-PTP: A Pedestrian Trajectory Prediction Method Based on Group Features
abstract
Group features have significant effects on pedestrian movement and constitute a focal point in pedestrian trajectory prediction research. In reality, pedestrians within a group exhibit notable consistency features due to their compact spatial positions, close destinations, and factors such as coordination within the group. In contrast, owing to the dispersed destinations among groups and the lack of coordination, there are significant differences in velocity and direction between the groups, leading to strong conflicts. However, existing pedestrian trajectory prediction models based on group features lack sufficient quantification of both within-group and between-group features. To address this problem, we propose Group-PTP, a novel pedestrian trajectory prediction model based on group features. Specifically, we first propose a group graph attention network-based group features aggregation method (Group-GAT). By quantifying and aggregating the intra-consistency and inter-conflict features exhibited by the groups, our method can better capture the features and interactions both within and between groups. Second, we propose a group multi-feature information representation model that fuses captured group aggregate features, pedestrian coordinates, surrounding pedestrian features, and obstacle features through fusion concatenation. Finally, we propose a multi-feature temporal convolutional network (MF-TCN) that embeds the impact weights of multi-feature information into pedestrian coordinates to obtain feature outputs and conducts temporal operations on feature outputs to predict future trajectories. The experimental results demonstrate that our proposed Group-PTP achieves state-of-the-art performance on several different trajectory prediction benchmarks.
Chuanyang Zhang, Guijuan Zhang, Zhuoran Zheng, Dianjie Lu
IEEE Trans. Multim.3
2024 Frequency Aware and Graph Fusion Network for Polyp Segmentation
abstract
Polyp segmentation plays a crucial role in the prevention of colon cancer. However, the diverse shapes of polyps and their similarity to normal areas in terms of color and texture make polyp segmentation a challenging task. Currently, most polyp segmentation methods solely focus on spatial domain features, ignoring the valuable features in the frequency domain. Consequently, many polyp segmentation algorithms struggle with the camouflage of polyps. To tackle this issue, we propose the Frequency Aware and Graph Fusion Network (FAGF-Net). Specifically, it begins with a Frequency-based Global Extraction Module (FGEM), which provides an initial estimation of the polyp regions to guide subsequent modules. Next, we design a Frequency-based Feature Attention Module (FFAM) that leverages amplitude and phase information to amplify appearance differences and enhance semantic representations. Moreover, we present a Graph-based Fusion Module (GFM), which infers the geometric characteristic of polyps through aggregating and interacting with enhanced features. Extensive experiments show that our method outperforms state-of-the-art methods with better quantitative and qualitative evaluations.
Yan Li 0196, Zhuoran Zheng, Wenqi Ren, Yunfeng Nie, Jingang Zhang, Xiuyi Jia
ICASSP2
2024 Rethinking Image Deraining via Text-guided Detail Reconstruction
abstract
Image deraining aims to recover clean images from degradation caused by rain streaks or raindrops of varying intensities. Recently many learning-based approaches have been proposed and achieved promising performance. However, these methods either focus on network architecture design or solely introduce image-level prior to the model. In this paper, we introduce text prior assisting the model in image deraining, as text descriptions have high flexibility and scalability. Text prior provides a wealth of semantic information to help the model achieve more detailed restoration, rather than blindly extrapolating details lost in the degraded image. To this end, we propose a novel image deraining framework based on the transformer, named TGDeraining. Specifically, to incorporate text prior into the framework, we design the Text Prior Embedded Transformer Block (TETB). TETB allows for dynamic guidance of the attention map guided by the text descriptions, thus emphasizing the restoration of critical missing details. The text prior is also fed to the feed-forward network to transform features in a controlled manner. Extensive experimental results demonstrate the effectiveness of our method in restoring a clear image using text as reference information.
Chen Wu 0006, Zhuoran Zheng, Pengwen Dai, Chenggang Shan, Xiuyi Jia
ICME2
2024 CLGNN: UAV Fault Diagnosis via Causal Learning and Graph Neural Network
abstract
To build an efficient unmanned aerial vehicle (UAV) fault diagnosis system, accurately modeling the relationship between sensor signals is a challenge. A common thought is to use a graph structure for modeling (Graph Neural Network), i.e., sensors as nodes and relationships between sensors are represented by edges. However, the Graph Neural Network (GNN) models, following the "Learning to Attend" principle, aim to maximize the mutual information between features and labels while minimizing training loss, without distinguishing the causal relationships between features and labels. The model’s lack of causal inference capability can lead to instability in model predictions, which ultimately shows up as a degradation in the model’s generalization performance on the test dataset. To address these issues, this paper proposes a GNN with causal learning to implement efficient UAV fault diagnosis, the model is called CLGNN. Specifically, the model first utilizes a GNN to model the sensor signals represented by the graph structure. Then, by intervening at the level of feature extraction, the non-causal components in the features are weakened to improve the model’s generalization performance in fault diagnosis. Extensive experimental results demonstrate that our method can robustly and accurately capture the anomalous signals of UAVs. In addition, we conduct a detailed analysis of the learned graphical representation of the sensor signals, exploring the decision-making basis of the model, which helps to boost the interpretability and reliability of the model.
Weiwei Li 0001, Zhuoran Zheng, Xiuyi Jia
IJCNN3
2024 Trusted re-weighting for label distribution learning
abstract
Label distribution learning (LDL) is a novel machine learning paradigm that aims to shift 0/1 labels into descriptive degrees to characterize the polysemy of instances. Since the description degree takes a value between 0 \ensuremath{\sim} 1, it is difficult for the annotator to accurately annotate each label. Therefore, the predictive ability of numerous LDL algorithms may be degraded by the presence of noise in the label space. To address this problem, we propose a novel stability-trust LDL framework that aims to reconstruct the feature space of an arbitrary LDL dataset by using feature decoupling and prototype guidance. Specifically, first, we use prototype learning to select reliable cluster centers (representative vectors of label distributions) to filter out a set of clean samples (with labeled noise) on the original dataset. Then, we decouple the feature space (eliminating correlations among features) by modeling a weight assigner that is learned on this clean sample set, thus assigning weights to each sample of the original dataset. Finally, all existing LDL algorithms can be trained on this new re-weighted dataset for the goal of robust modeling. In addition, we create a new image dataset to support the training and testing of compared models. Experimental results demonstrate that the proposed framework boosts the performance of the LDL algorithm on datasets with label noise.
Zhuoran Zheng, Chen Wu 0006, Yeying Jin, Xiuyi Jia
UAI1
2024 Single UHD image dehazing via Interpretable Pyramid Network
Boxue Xiao, Zhuoran Zheng, Yunliang Zhuang, Chen Lyu 0001, Xiuyi Jia
Signal Process.2
2024 Polyp-DAM: Polyp Segmentation via Depth Anything Model
abstract
Recently, large models (Segment Anything model) came on the scene to provide a new baseline for polyp segmentation tasks. This demonstrates that large models with a sufficient image level prior can achieve promising performance on a given task. In this paper, we unfold a new perspective on polyp segmentation modeling by leveraging the Depth Anything Model (DAM) to provide depth prior to polyp segmentation models. Specifically, the input polyp image is first passed through a frozen DAM to generate a depth map. The depth map and the input polyp images are then concatenated and fed into a convolutional neural network with multiscale to generate segmented images. Extensive experimental results demonstrate the effectiveness of our method, and in addition, we observe that our method still performs well on images of polyps with noise.
Zhuoran Zheng, Chen Wu 0006, Yeying Jin, Xiuyi Jia
IEEE Signal Process. Lett.1
2024 Dimensional Transformation Mixer for Ultra-High-Definition Industrial Camera Dehazing
abstract
Haze severely affects the reliability of vision-based industrial systems, and most of the current dehazing methods are not applicable to images captured by ultra-high-definition (UHD) industrial cameras. In this article, we propose a novel dimensional transformation mixer (DMixer) model for recovering haze-free images from UHD haze images. In DMixer, the dimensional transformation module encodes the complete image in multiple stages from different perspectives and associates features from different views by permuting the tensor for efficient long-range dependency modeling. In this way, the global perception capabilities of DMixer are complemented, allowing the quality of the reconstructed UHD images to be improved. Furthermore, DMixer employs a dual-stream network framework that combines local and multiscale features, allowing DMixer to better trade off the performance and efficiency for real industrial systems. Our proposed method enables the real-time processing of UHD images ($\sim$59 fps). Extensive results show that the proposed model outperforms current state-of-the-art methods, with PSNR improvements of 1.94 and 0.95 dB on the 4KID and O-HAZE datasets, respectively.
Yunliang Zhuang, Zhuoran Zheng, Lei Lyu 0001, Xiuyi Jia, Chen Lyu 0001
IEEE Trans. Ind. Informatics2
2023 Image Restoration via UAVFormer for Under-Display Camera of UAV
abstract
The exposed cameras of UAVs can shake, shift, or even malfunction under the influence of harsh weather, while the add-on devices (Dupont lines) are very vulnerable to dam-age. Although we can place a low-cost transparent film overlay around the camera to protect it, this would also introduce image degradation issues (such as oversaturation, astigmatism, etc). To tackle the image degradation problem caused by overlaying transparent film, in this paper we propose a novel method to enhance the visual experience by adapting a deep network with UAV characteristics. Specifically, we propose a customized Transformer named UAVFormer to recover the image, which has a key module at each stage based on the Swin Transformer with local awareness (LAT). In the end, we use an evidential fusion algorithm to integrate the generated images at each stage to obtain a high-quality result. Furthermore, we create a high-resolution under-display camera dataset to support the training and testing of compared models. Our model can conduct high-quality recovery of images of 2K resolution on some embedded devices (Raspberry Pi 4b) in realtime. The URL for the code at https://github.com/zzr-idam/UAVFormer.
Zhuoran Zheng, Xiuyi Jia
IROS1
2023 Dual-Domain Learning Network for Polyp Segmentation
Yan Li 0196, Zhuoran Zheng, Wenqi Ren, Yunfeng Nie, Jingang Zhang, Xiuyi Jia
IWDW2
2022 UHD Underwater Image Enhancement via Frequency-Spatial Domain Aware Network
Yiwen Wei, Zhuoran Zheng, Xiuyi Jia
ACCV (3)2
2022 Unpaired Deep Image Dehazing Using Contrastive Disentanglement Learning
Xiang Chen 0015, Zhentao Fan, Pengpeng Li 0001, Longgang Dai, Caihua Kong, Zhuoran Zheng, Yufeng Li 0001
ECCV (17)6
2022 HELoC: hierarchical contrastive learning of source code representation
abstract
Abstract syntax trees (ASTs) play a crucial role in source code representation. However, due to the large number of nodes in an AST and the typically deep AST hierarchy, it is challenging to learn the hierarchical structure of an AST effectively. In this paper, we propose HELoC, a hierarchical contrastive learning model for source code representation. To effectively learn the AST hierarchy, we use contrastive learning to allow the network to predict the AST node level and learn the hierarchical relationships between nodes in a self-supervised manner, which makes the representation vectors of nodes with greater differences in AST levels farther apart in the embedding space. By using such vectors, the structural similarities between code snippets can be measured more precisely. In the learning process, a novel GNN (called Residual Self-attention Graph Neural Network, RSGNN) is designed, which enables HELoC to focus on embedding the local structure of an AST while capturing its overall structure. HELoC is self-supervised and can be applied to many source code related downstream tasks such as code classification, code clone detection, and code clustering after pre-training. Our extensive experiments demonstrate that HELoC outperforms the state-of-the-art source code representation models.
Hongyu Zhang 0002, Chen Lyu 0001, Zhuoran Zheng, Lei Lyu 0001, Songlin Hu 0001
ICPC6
2022 UHD Low-light image enhancement via interpretable bilateral learning
Qiaowanni Lin, Zhuoran Zheng, Xiuyi Jia
Inf. Sci.2
2021 Ultra-High-Definition Image Dehazing via Multi-Guided Bilateral Learning
abstract
Convolutional neural networks (CNNs) have achieved significant success in the single image dehazing task. Unfortunately, most existing deep dehazing models have high computational complexity, which hinders their application to high-resolution images, especially for UHD (ultra-high-definition) or 4K resolution images. To address the problem, we propose a novel network capable of real-time dehazing of 4K images on a single GPU, which consists of three deep CNNs. The first CNN extracts haze-relevant features at a reduced resolution of the hazy input and then fits locally-affine models in the bilateral space. Another CNN is used to learn multiple full-resolution guidance maps corresponding to the learned bilateral model. As a result, the feature maps with high-frequency can be reconstructed by multi-guided bilateral upsampling. Finally, the third CNN fuses the high-quality feature maps into a dehazed image. In addition, we create a large-scale 4K image dehazing dataset to support the training and testing of compared models. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art dehazing approaches on various benchmarks.
Zhuoran Zheng, Wenqi Ren, Xiaochun Cao, Xiaobin Hu, Tao Wang 0053, Fenglong Song, Xiuyi Jia
CVPR1
2021 Ultra-High-Definition Image HDR Reconstruction via Collaborative Bilateral Learning
abstract
Existing single image high dynamic range (HDR) reconstruction methods attempt to expand the range of illuminance. They are not effective in generating plausible textures and colors in the reconstructed results, especially for high-density pixels in ultra-high-definition (UHD) images. To address these problems, we propose a new HDR reconstruction network for UHD images by collaboratively learning color and texture details. First, we propose a dual-path network to extract the content and chromatic features at a reduced resolution of the low dynamic range (LDR) input. These two types of features are used to fit bilateral-space affine models for real-time HDR reconstruction. To extract the main data structure of the LDR input, we propose to use 3D Tucker decomposition and reconstruction to prevent pseudo edges and noise amplification in the learned bilateral grid. As a result, the high-quality content and chromatic features can be reconstructed capitalized on guided bilateral upsampling. Finally, we fuse these two full-resolution feature maps into the HDR reconstructed results. Our proposed method can achieve real-time processing for UHD images (about 160 fps). Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art HDR reconstruction approaches on public benchmarks and real-world UHD images.
Zhuoran Zheng, Wenqi Ren, Xiaochun Cao, Tao Wang 0053, Xiuyi Jia
ICCV1
2021 Saliency Detection Framework Based on Deep Enhanced Attention Network
Xing Sheng 0001, Zhuoran Zheng, Chunmeng Kang, Yunliang Zhuang, Lei Lyu 0001, Chen Lyu 0001
ICONIP (4)2
2021 Code Representation Based on Hybrid Graph Modelling
Zhuoran Zheng, Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001
ICONIP (5)3
2021 TreeBERT: A tree-based pre-trained model for programming language
abstract
Source code can be parsed into the abstract syntax tree (AST) based on defined syntax rules. However, in pre-training, little work has considered the incorporation of tree structure into the learning process. In this paper, we present TreeBERT, a tree-based pre-trained model for improving programming language-oriented generation tasks. To utilize tree structure, TreeBERT represents the AST corresponding to the code as a set of composition paths and introduces node position embedding. The model is trained by tree masked language modeling (TMLM) and node order prediction (NOP) with a hybrid objective. TMLM uses a novel masking strategy designed according to the tree’s characteristics to help the model understand the AST and infer the missing semantics of the AST. With NOP, TreeBERT extracts the syntactical structure by learning the order constraints of nodes in AST. We pre-trained TreeBERT on datasets covering multiple programming languages. On code summarization and code documentation tasks, TreeBERT outperforms other pre-trained models and state-of-the-art models designed for these tasks. Furthermore, TreeBERT performs well when transferred to the pre-trained unseen programming language.
Zhuoran Zheng, Chen Lyu 0001, Lei Lyu 0001
UAI2