Shi-Chen Zhang

dblp:351/1012 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0007-1973-306XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 53% Representation and self-supervised learning · 16% Deep learning architectures and training · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.722025
Low-Resolution Self-Attention for Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment · ICCV 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
efficient semantic segmentation
0.912025
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment · ICCV 2025
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.912025
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment · ICCV 2025
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.912025
Low-Resolution Self-Attention for Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › medical image analysis
medical image detection
0.812024
Revisiting Computer-Aided Tuberculosis Diagnosis · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Medical and health informatics
computer-aided diagnosis
0.812024
Revisiting Computer-Aided Tuberculosis Diagnosis · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Medical and health informatics › clinical diagnosis
tuberculosis diagnosis
0.812024
Revisiting Computer-Aided Tuberculosis Diagnosis · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Segmentation and scene understanding
medical image segmentation
0.212024
Revisiting Computer-Aided Tuberculosis Diagnosis · IEEE Trans. Pattern Anal. Mach. Intell. 2024

Methods — techniques the papers use, named apart from their topics

symmetric search attention · 1.5symmetric positional encoding · 1.5offset learning · 0.9encoder-decoder architecture · 0.9dual-branch network · 0.9depthwise convolution · 0.9
YearPublicationVenuePosition
2025 Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
abstract
Semantic segmentation is fundamental to vision systems requiring pixel-level scene understanding, yet deploying it on resource-constrained devices demands efficient architectures. Although existing methods achieve real-time inference through lightweight designs, we reveal their inherent limitation: misalignment between class representations and image features caused by a per-pixel classification paradigm. With experimental analysis, we find that this paradigm results in a highly challenging assumption for efficient scenarios: Image pixel features should not vary for the same category in different images. To address this dilemma, we propose a coupled dual-branch offset learning paradigm that explicitly learns feature and class offsets to dynamically refine both class representations and spatial image features. Based on the proposed paradigm, we construct an efficient semantic segmentation network, OffSeg. Notably, the offset learning paradigm can be adopted to existing methods with no additional architectural changes. Extensive experiments on four datasets, including ADE20K, Cityscapes, COCO-Stuff-164K, and Pascal Context, demonstrate consistent improvements with negligible parameters. For instance, on the ADE20K dataset, our proposed offset learning paradigm improves SegFormer-B0, SegNeXt-T, and Mask2Former-Tiny by 2.7%, 1.9%, and 2.6% mIoU, respectively, with only 0.1-0.2M additional parameters required.
Shi-Chen Zhang, Yu-Huan Wu, Qibin Hou, Ming-Ming Cheng
ICCV1
2025 Low-Resolution Self-Attention for Semantic Segmentation
abstract
Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space, regardless of the input image's resolution, with additional $\text{3}\times \text{3}$3×3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20 K, COCO-Stuff, and CityScapes datasets demonstrate that LRFormer outperforms state-of-the-art models.
Yu-Huan Wu, Shi-Chen Zhang, Yun Liu 0011, Le Zhang 0001, Xin Zhan, Daquan Zhou, Jiashi Feng, Ming-Ming Cheng, Liangli Zhen
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Revisiting Computer-Aided Tuberculosis Diagnosis
abstract
Tuberculosis (TB) is a major global health threat, causing millions of deaths annually. Although early diagnosis and treatment can greatly improve the chances of survival, it remains a major challenge, especially in developing countries. Recently, computer-aided tuberculosis diagnosis (CTD) using deep learning has shown promise, but progress is hindered by limited training data. To address this, we establish a large-scale dataset, namely the Tuberculosis X-ray (TBX11 K) dataset, which contains 11 200 chest X-ray (CXR) images with corresponding bounding box annotations for TB areas. This dataset enables the training of sophisticated detectors for high-quality CTD. Furthermore, we propose a strong baseline, SymFormer, for simultaneous CXR image classification and TB infection area detection. SymFormer incorporates Symmetric Search Attention (SymAttention) to tackle the bilateral symmetry property of CXR images for learning discriminative features. Since CXR images may not strictly adhere to the bilateral symmetry property, we also propose Symmetric Positional Encoding (SPE) to facilitate SymAttention through feature recalibration. To promote future research on CTD, we build a benchmark by introducing evaluation metrics, evaluating baseline models reformed from existing detectors, and running an online challenge. Experiments show that SymFormer achieves state-of-the-art performance on the TBX11 K dataset.
Yun Liu 0011, Yu-Huan Wu, Shi-Chen Zhang, Li Liu 0002, Min Wu 0008, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.3