EDBT 2026 Demo / reviewers in the wild / expert
Shengxuming Zhang
dblp:355/1122
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-8827-9012ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Attention Patterns in Vision Transformers for Robustness
Haofei Zhang, Hanyang Yuan, Haoze Jiang, Jiacong Hu, Shengxuming Zhang, Mingli Song |
ICIC (7) | 9 |
| 2025 | Dataset Ownership Verification for Pre-Trained Masked ModelsabstractHigh-quality open-source datasets have emerged as a pivotal catalyst driving the swift advancement of deep learning, while facing the looming threat of potential exploitation. Protecting these datasets is of paramount importance for the interests of their owners. The verification of dataset ownership has evolved into a crucial approach in this domain; however, existing verification techniques are predominantly tailored to supervised models and contrastive pre-trained models, rendering them ill-suited for direct application to the increasingly prevalent masked models. In this work, we introduce the inaugural methodology addressing this critical, yet unresolved challenge, termed Dataset Ownership Verification for Masked Modeling (DOV4MM). The central objective is to ascertain whether a suspicious black-box model has been pre-trained on a particular unlabeled dataset, thereby assisting dataset owners in safeguarding their rights. DOV4MM is grounded in our empirical observation that when a model is pre-trained on the target dataset, the difficulty of reconstructing masked information within the embedding space exhibits a marked contrast to models not pre-trained on that dataset. We validated the efficacy of DOV4MM through ten masked image models on ImageNet-1K and four masked language models on WikiText-103. The results demonstrate that DOV4MM rejects the null hypothesis, with a $p$-value considerably below 0.05, surpassing all prior approaches. Code is available at https://github.com/xieyc99/DOV4MM. Yuechen Xie, Jie Song 0011, Yicheng Shan, Yuanyu Wan, Shengxuming Zhang, Jiarui Duan, Mingli Song |
ICCV | 6 |
| 2025 | Large Vision-Language Models are Generalist Solvers For Pathology TasksabstractLeveraging the powerful capabilities of large language models (LLMs), large vision-language models (LVLMs) can perform a wide variety of tasks based on input images and user instructions. However, existing pathology-focused LVLMs are limited to relatively simple tasks, such as image captioning, visual question answering, and generating brief pathology reports, which restricts their clinical applicability. To enhance the practicality of pathology LVLMs and explore their performance boundaries across pathology tasks, we curated a multi-task pathology instruction-following dataset that better aligns with clinical needs. This dataset encompasses tasks such as cancer classification and grading, molecular subtype identification, and the detection of structures like nuclei, blood vessels, nerves, and lymph nodes. Extensive experiments were conducted on this dataset to identify key factors influencing the performance of LVLMs on these pathology tasks, and optimal solutions were proposed. Our findings provide valuable insights to advance the clinical application of large vision-language models in pathology. Shengxuming Zhang, Hengrui Lou, Jing Zhang 0120, Xiuming Zhang, Mingli Song, Zunlei Feng |
ICIP | 1 |
| 2025 | L-Diffusion: Laplace Diffusion for Efficient Pathology Image SegmentationabstractPathology image segmentation plays a pivotal role in artificial digital pathology diagnosis and treatment. Existing approaches to pathology image segmentation are hindered by labor-intensive annotation processes and limited accuracy in tail-class identification, primarily due to the long-tail distribution inherent in gigapixel pathology images. In this work, we introduce the Laplace Diffusion Model, referred to as L-Diffusion, an innovative framework tailored for efficient pathology image segmentation. L-Diffusion utilizes multiple Laplace distributions, as opposed to Gaussian distributions, to model distinct components—a methodology supported by theoretical analysis that significantly enhances the decomposition of features within the feature space. A sequence of feature maps is initially generated through a series of diffusion steps. Following this, contrastive learning is employed to refine the pixel-wise vectors derived from the feature map sequence. By utilizing these highly discriminative pixel-wise vectors, the segmentation module achieves a harmonious balance of precision and robustness with remarkable efficiency. Extensive experimental evaluations demonstrate that L-Diffusion attains improvements of up to 7.16%, 26.74%, 16.52%, and 3.55% on tissue segmentation datasets, and 20.09%, 10.67%, 14.42%, and 10.41% on cell segmentation datasets, as quantified by DICE, MPA, mIoU, and FwIoU metrics. The source are available at https://github.com/Lweihan/LDiffusion. Linyun Zhou, Yang Jian, Shengxuming Zhang, Xiangtong Du, Xiuming Zhang, Jing Zhang 0120, Chaoqing Xu, Mingli Song, Zunlei Feng |
ICML | 4 |
| 2025 | Self-calibration Enhanced Whole Slide Pathology Image AnalysisabstractPathology images are considered the ``gold standard" for cancer diagnosis and treatment, with gigapixel images providing extensive tissue and cellular information. Existing methods fail to simultaneously extract global structural and local detail features for comprehensive pathology image analysis efficiently. To address these limitations, we propose a self-calibration enhanced framework for whole slide pathology image analysis, comprising three components: a global branch, a focus predictor, and a detailed branch. The global branch initially classifies using the pathological thumbnail, while the focus predictor identifies relevant regions for classification based on the last layer features of the global branch. The detailed extraction branch then assesses whether the magnified regions correspond to the lesion area. Finally, a feature consistency constraint between the global and detail branches ensures that the global branch focuses on the appropriate region and extracts sufficient discriminative features for final identification. These focused discriminative features can facilitate the discovery of novel prognostic tumor markers, from the perspective of feature uniqueness and tissue spatial distribution. Extensive experiment results demonstrate that the proposed framework can rapidly deliver accurate and explainable results for pathological grading and prognosis tasks. Haoming Luo, Xiaotian Yu, Shengxuming Zhang, Jiabin Xia, Jian Yang 0003, Yuning Sun, Xiuming Zhang, Jing Zhang 0120, Zunlei Feng |
IJCAI | 3 |
| 2025 | DenseSAM: Semantic Enhance SAM for Efficient Dense Object SegmentationabstractDense object segmentation is essential for various applications, particularly in pathology image and remote sensing image analysis. However, distinguishing numerous similar and densely packed objects in this task presents significant challenges. Several methods, including CNN- and ViT-based approaches, have been proposed to tackle these issues. Yet, models trained on limited datasets exhibit limited generalization ability. The Segment Anything Model (SAM) has recently achieved significant progress in zero-shot segmentation but relies heavily on precise positional guidance. However, providing numerous accurate location prompts in dense scenarios is time-consuming. To overcome this limitation, we conducted an in-depth exploration of the SAM mechanism and found that its strong generalization ability stems from the encoder’s edge detection capability, which is semantically independent, making location prompts essential for segmentation. This insight inspired the development of DenseSAM, which replaces location prompts with semantic guidance for automatic segmentation in dense scenarios. Specifically, it uses local details to weaken the edges of background objects, leverages global context to enhance intra-class feature similarity, while further increasing contrast with the background, and integrates a dual-head decoding process to enable lightweight automatic semantic segmentation. Extensive experiments on pathology images demonstrate that DenseSAM delivers remarkable performance with minimal training parameters, providing a cost-effective and efficient solution. Moreover, experiments on remote sensing images further validate its excellent scalability, making DenseSAM suitable for various dense object segmentation domains. The code is available at https://github.com/imAzhou/DenseSAM. Linyun Zhou, Jiacong Hu, Shengxuming Zhang, Xiangtong Du, Mingli Song, Xiuming Zhang, Zunlei Feng |
IJCAI | 3 |
| 2024 | E3V-K5: An Authentic Benchmark for Redefining Video-Based Energy Expenditure Estimation
Shengxuming Zhang, Xinyu Wang 0001, Zunlei Feng, Mingli Song |
ECCV (35) | 1 |
| 2024 | Target Optimization Direction Guided Transfer Learning for Image ClassificationabstractAt present, deep learning has made impressive achievements in various fields; however, effectively training deep neural networks on small data sets remains a significant challenge. Transfer learning, as a method of efficient training across multiple tasks, has been widely used to solve this problem. However, when the domain gap or the data volume difference between the two tasks is too large, the transfer learning may not perform well, and other optimization methods will be required to improve the performance. In this paper, we propose a new transfer learning method guided by the direction of objective optimization from the perspective of gradient. This method guides the gradient direction of the source task towards the gradient direction of the target task. In several similar and conflicting tasks, this method has achieved good results in efficiency and performance. In comparison with other transfer learning methods, the results shown by this method are generally better. Kelvin Ting Zuo Han, Shengxuming Zhang, Gerard Marcos Freixas, Zunlei Feng, Cheng Jin 0001 |
ICASSP | 2 |
| 2024 | Loose Lesion Location Self-supervision Enhanced Colorectal Cancer Diagnosis
Tianhong Gao, Jie Song 0011, Xiaotian Yu, Shengxuming Zhang, Xiuming Zhang, Zipeng Zhong, Mingli Song, Zunlei Feng |
MICCAI (11) | 4 |
| 2024 | Vision Mamba MenderabstractMamba, a state-space model with selective mechanisms and hardware-aware architecture, has demonstrated outstanding performance in long sequence modeling tasks, particularly garnering widespread exploration and application in the field of computer vision. While existing works have mixed opinions of its application in visual tasks, the exploration of its internal workings and the optimization of its performance remain urgent and worthy research questions given its status as a novel model. Existing optimizations of the Mamba model, especially when applied in the visual domain, have primarily relied on predefined methods such as improving scanning mechanisms or integrating other architectures, often requiring strong priors and extensive trial and error. In contrast to these approaches, this paper proposes the Vision Mamba Mender, a systematic approach for understanding the workings of Mamba, identifying flaws within, and subsequently optimizing model performance. Specifically, we present methods for predictive correlation analysis of Mamba's hidden states from both internal and external perspectives, along with corresponding definitions of correlation scores, aimed at understanding the workings of Mamba in visual recognition tasks and identifying flaws therein. Additionally, tailored repair methods are proposed for identified external and internal state flaws to eliminate them and optimize model performance. Extensive experiments validate the efficacy of the proposed methods on prevalent Mamba architectures, significantly enhancing Mamba's performance. Jiacong Hu, Anda Cao, Zunlei Feng, Shengxuming Zhang, Lingxiang Jia, Mingli Song |
NeurIPS | 4 |
| 2024 | Noise is the fatal poison: A Noise-aware Network for noisy dataset classification
Xiaotian Yu, Shengxuming Zhang, Lingxiang Jia, Mingli Song, Zunlei Feng |
Neurocomputing | 2 |
| 2024 | PatchDetector: Pluggable and non-intrusive patch for small object detection
Linyun Zhou, Shengxuming Zhang, Zunlei Feng, Mingli Song |
Neurocomputing | 2 |
| 2023 | A Loopback Network for Explainable Microvascular Invasion ClassificationabstractMicrovascular invasion (MVI) is a critical factor for prognosis evaluation and cancer treatment. The current diagnosis of MVI relies on pathologists to manually find out cancerous cells from hundreds of blood vessels, which is time-consuming, tedious, and subjective. Recently, deep learning has achieved promising results in medical image analysis tasks. However, the unexplainability of black box models and the requirement of massive annotated samples limit the clinical application of deep learning based diagnostic methods. In this paper, aiming to develop an accurate, objective, and explainable diagnosis tool for MVI, we propose a Loopback Network (LoopNet) for classifying MVI efficiently. With the image-level category annotations of the collected Pathologic Vessel Image Dataset (PVID), LoopNet is devised to be composed binary classification branch and cell locating branch. The latter is devised to locate the area of cancerous cells, regular non-cancerous cells, and background. For healthy samples, the pseudo masks of cells supervise the cell locating branch to distinguish the area of regular non-cancerous cells and background. For each MVI sample, the cell locating branch predicts the mask of cancerous cells. Then the masked cancerous and non-cancerous areas of the same sample are input back to the binary classification branch separately. The loopback between two branches enables the category label to supervise the cell locating branch to learn the locating ability for cancerous areas. Experiment results show that the proposed LoopNet achieves 97.5% accuracy on MVI classification. Surprisingly, the proposed loopback mechanism not only enables LoopNet to predict the cancerous area but also facilitates the classification backbone to achieve better classification performance. Shengxuming Zhang, Tianqi Shi, Xiuming Zhang, Jie Lei 0002, Zunlei Feng, Mingli Song |
CVPR | 1 |