Shun Zou

dblp:368/3853 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking Block Boundaries: Anchor-based History-stable Decoding for Diffusion Large Language Models
abstract
Shun Zou, Yong Wang, Zehui Chen, Lin Chen, Chongyang Tao, Feng Zhao, Xiangxiang Chu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shun Zou, Chongyang Tao, Feng Zhao 0004, Xiangxiang Chu
ACL (1)1
2026 BiMarker: Enhancing text watermark detection for large language models with bipolar watermarks
Qiuping Yi, Zongcheng Ji, Yijian Lu, Shun Zou, Yanqi Li, Keyang Xiao, Hongliang Liang
Neurocomputing5
2026 BinSleuth: Scalable vulnerability discovery in binaries via automated source identification and path optimization
Shun Zou, Hongliang Liang
J. Syst. Softw.2
2025 OCTAMamba: A State-Space Model Approach for Precision OCTA Vasculature Segmentation
abstract
Optical Coherence Tomography Angiography (OCTA) is a crucial imaging technique for visualizing retinal vasculature and diagnosing eye diseases such as diabetic retinopathy and glaucoma. However, precise segmentation of OCTA vasculature remains challenging due to the multi-scale vessel structures and noise from poor image quality and eye lesions. In this study, we proposed OCTAMamba, a novel U-shaped network based on the Mamba architecture, designed to segment vasculature in OCTA accurately. OCTAMamba integrates a Quad Stream Efficient Mining Embedding Module for local feature extraction, a Multi-Scale Dilated Asymmetric Convolution Module to capture multi-scale vasculature, and a Focused Feature Recalibration Module to filter noise and highlight target areas. Our method achieves efficient global modeling and local feature extraction while maintaining linear complexity, making it suitable for low-computation medical applications. Extensive experiments on the OCTA 3M, OCTA 6M, and ROSSA datasets demonstrated that OCTAMamba outperforms state-of-the-art methods, providing a new reference for efficient OCTA segmentation. Code is available at https://github.com/zs1314/OCTAMamba
Shun Zou, Zhuo Zhang 0007, Guangwei Gao
ICASSP1
2025 BMC-Net: A Framework for IDH Genotyping of Gliomas Based on Bi-Directional Mamba Sequences
Shuaidan Wang, Shun Zou, Yuhan He, Zhuo Zhang 0020
ICIC (28)2
2025 MambaMIC: An Efficient Baseline for Microscopic Image Classification with State Space Models
abstract
In recent years, CNN and Transformer-based methods have made significant progress in Microscopic Image Classification (MIC). However, existing approaches still face the dilemma between global modeling and efficient computation. While the Selective State Space Model (SSM) can simulate long-range dependencies with linear complexity, it still encounters challenges in MIC, such as local pixel forgetting, channel redundancy, and lack of local perception. To address these issues, we propose a simple yet efficient vision backbone for MIC tasks, named MambaMIC. Specifically, we introduce a Local-Global dual-branch aggregation module: the MambaMIC Block, designed to effectively capture and fuse local connectivity and global dependencies. In the local branch, we use local convolutions to capture pixel similarity, mitigating local pixel forgetting and enhancing perception. In the global branch, SSM extracts global dependencies, while Locally Aware Enhanced Filter reduces channel redundancy and local pixel forgetting. Additionally, we design a Feature Modulation Interaction Aggregation Module for deep feature interaction and key feature re-localization. Extensive benchmarking shows that MambaMIC achieves state-of-the-art performance across five datasets. code is available at https://zs1314.github.io/MambaMIC.
Shun Zou, Zhuo Zhang 0007, Guangwei Gao
ICME1
2025 Fraesormer: Learning Adaptive Sparse Transformer for Efficient Food Recognition
abstract
In recent years, Transformer has witnessed significant progress in food recognition. However, most existing approaches still face two critical challenges in lightweight food recognition: (1) the quadratic complexity and redundant feature representation from interactions with irrelevant tokens; (2) static feature recognition and single-scale representation, which overlook the unstructured, non-fixed nature of food images and the need for multi-scale features. To address these, we propose an adaptive and efficient sparse Transformer architecture (Fraesormer) with two core designs: Adaptive Top-k Sparse Partial Attention (ATK-SPA) and Hierarchical Scale-Sensitive Feature Gating Network (HSSFGN). ATK-SPA uses a learnable Gated Dynamic Top-K Operator (GDTKO) to retain critical attention scores, filtering low query-key matches that hinder feature aggregation. It also introduces a partial channel mechanism to reduce redundancy and promote expert information flow, enabling local-global collaborative modeling. HSSFGN employs gating mechanism to achieve multi-scale feature representation, enhancing contextual semantic information. Extensive experiments show that Fraesormer outperforms state-of-the-art methods. code is available at https://zs1314.github.io/Fraesormer.
Shun Zou, Mingya Zhang, Shipeng Luo, Zhihao Chen 0014, Guangwei Gao
ICME1
2025 Learning Dual-Domain Multi-Scale Representations for Single Image Deraining
abstract
Existing image deraining methods typically rely on single-input, single-output, and single-scale architectures, which overlook the joint multi-scale information between external and internal features. Furthermore, single-domain representations are often too restrictive, limiting their ability to handle the complexities of real-world rain scenarios. To address these challenges, we propose a novel Dual-Domain Multi-Scale Representation Network (DMSR). The key idea is to exploit joint multi-scale representations from both external and internal domains in parallel while leveraging the strengths of both spatial and frequency domains to capture more comprehensive properties. Specifically, our method consists of two main components: the Multi-Scale Progressive Spatial Refinement Module (MPSRM) and the Frequency Domain Scale Mixer (FDSM). The MPSRM enables the interaction and coupling of multi-scale expert information within the internal domain using a hierarchical modulation and fusion strategy. The FDSM extracts multi-scale local information in the spatial domain, while also modeling global dependencies in the frequency domain. Extensive experiments show that our model achieves state-of-the-art performance across six benchmark datasets.
Shun Zou, Mingya Zhang, Shipeng Luo, Guangwei Gao, Guo-Jun Qi
ICME1
2025 Cross Paradigm Representation and Alignment Transformer for Image Deraining
abstract
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications.
Shun Zou, Juncheng Li 0003, Guangwei Gao, Guo-Jun Qi
ACM Multimedia1
2024 Light Fields Stitching for Windowed-6DoF VR Content
abstract
Windowed six degrees of Freedom (Windowed-6DoF) virtual reality (VR) content that provides users an immersive feeling of walking through a 3D 360 VR space with constrained rotational movements around X and Y axes and constrained translational movements along Z axis is important for the development of VR. To facilitate this windowed-6DoF immersive feeling, light fields (LFs) from multiple perspectives within the windowed 6-DoF space are captured. In contrast to employing a large-scale camera array for LF capture, utilizing hand-held plenoptic cameras offers a more portable and versatile solution, thereby promoting practical applications. However, how to stitch the LFs at different rotational angles containing motion parallax is challenging. In this paper, a novel LF stitching method is proposed to generate windowed-6DoF LFs. First, multi-concentric spherical modeling is proposed to parameterize the recorded LFs to eliminate projection biases in the registration process. Then, a global-local adaptive LF registration is proposed by developing incremental multi-layer global-local adaptive homographies based on the 4D light field feature (LiFF), incremental strategy and depth layer maps (DLMs) to eliminate parallax errors. Testing on the LFs captured in both indoor and outdoor scenes with different focal lengths, quantities of LFs and scales of translation and rotation, the proposed method outperforms the existing approaches in terms of subjective quality, objective quality, light field consistency and content production robustness, which can produce VR content of superior quality more reliably.
Yihui Fan, Xin Jin 0002, Siyao Zhou 0003, Shun Zou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Underwater Refractive Stereo Vision Using Ray Tracing for Dome Port Camera
abstract
Stereo vision is one of the most important ways to explore the underwater world. However, the refraction occurs on the dome port camera housing, resulting in distorted images and errors in the computed 3D information. In this paper, we proposed an underwater refractive stereo vision method that can calibrate the camera’s parameters and obtain 3D information underwater. First, a dome port refractive ray tracing model is proposed to describe the projection relation between the object point and the image point. Then, a refractive calibration method is proposed to calibrate the camera’s intrinsic and extrinsic parameters without pre-calibration in air. Finally, a refractive rectification method is proposed for disparity search along polar lines and calculating 3D information. Experimental results show that the proposed method outperforms existing methods in both simulation and real imaging scenarios.
Weijin Lv, Jiang Guotai, Shun Zou, Yihui Fan
VCIP4
2023 Multi-View Image Rectification for UAV-Captured Image Sequences
abstract
The image sequences captured by Unmanned Aerial Vehicles (UAVs) can be applied to many computer vision tasks. However, due to the instability of UAV flight, the captured image sequences will deviate from the preset trajectory and pose, which reduce the quality of subsequent applications such as panoramic image stitching. In this paper, a novel method is proposed to rectify UAV-captured image sequences by transforming the images to a regular trajectory with the uniform pose. First, to minimize the total transformation deviation, virtual regular camera trajectory is derived by minimizing the global error of coordinates between actual and virtual camera trajectories. Then, camera-pose-relevant local homography is proposed by inserting the camera pose into local homography to transform the images to the derived virtual trajectory with the uniform pose and correct translation parallax. The experimental results demonstrate the effectiveness of the proposed rectification algorithm from both theoretical and application levels.
Shun Zou, Yihui Fan
VCIP1