VLDB 2026 Research / reviewers in the wild / expert
Sheng Zhang 0024
dblp:69/6137-24
· DBLP profile ↗
16ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-0516-6491ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test-time generative augmentation for medical image segmentation
Xiao Ma 0011, Yuhui Tao, Zetian Zhang, Yuhan Zhang 0001, Xi Wang 0013, Sheng Zhang 0024, Zexuan Ji, Yizhe Zhang 0001, Qiang Chen 0004, Guang Yang 0006 |
Medical Image Anal. | 6 |
| 2026 | Dynamical multi-order responses and global semantic-infused adversarial learning: A robust airway segmentation methodabstractAutomated airway segmentation in computerized tomography (CT) images is crucial for the accurate diagnosis of lung diseases. However, the scarcity of manual annotations hinders the efficacy of supervised learning, while unconstrained intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order responses and Global Semantic-infused Adversarial network (DMGSA), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to empower the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles; (3) we introduce the Adversarial Learning (AL) on the top of MONR module to discern nuances between real and fake images, focusing on capturing the textural features of terminal bronchioles. For the supervised branch, we propose an innovative Generalized Mean pooling based Global Semantic-infused (GMGS) module to ulteriorly improve the robustness. Ultimately, we have verified the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly. Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Yongkai Liu, Giorgos Papanastasiou, Zhifan Gao, Shuo Li 0001, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 1 |
| 2025 | DMRN: A Dynamical Multi-Order Response Network for the Robust Lung Airway SegmentationabstractAutomated airway segmentation in CT images is crucial for lung diseases' diagnosis. However, manual annotation scarcity hinders supervised learning efficacy, while unlimited intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order Response Network (DMRN), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to enable the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles. For the supervised branch, we directly predict the final full segmentation map by the large-ratio cube-masked input instead of full input. Ultimately, we verify the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly. Code will be released in the future. Sheng Zhang 0024, Jinge Wu, Junzhi Ning, Guang Yang 0006 |
WACV | 1 |
| 2025 | A lung structure and function information-guided residual diffusion model for predicting idiopathic pulmonary fibrosis progression
Caiwen Jiang, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Simon Walsh, Guang Yang 0006, Dinggang Shen |
Medical Image Anal. | 5 |
| 2025 | Unpaired translation of chest X-ray images for lung opacity diagnosis via adaptive activation masks and cross-domain alignmentabstractChest X-ray radiographs (CXRs) play a pivotal role in diagnosing and monitoring cardiopulmonary diseases. However, lung opacities in CXRs frequently obscure anatomical structures, impeding clear identification of lung borders and complicating localisation of pathology. This challenge significantly hampers segmentation accuracy and precise lesion identification, crucial for diagnosis. To tackle these issues, our study proposes an unpaired CXR translation framework that converts CXRs with lung opacities into counterparts without lung opacities while preserving semantic features. Central to our approach is the use of adaptive activation masks to selectively modify opacity regions in lung CXRs. Cross-domain alignment ensures translated CXRs without opacity issues align with feature maps and prediction labels from a pre-trained CXR lesion classifier, facilitating the interpretability of the translation process. We validate our method using RSNA, MIMIC-CXR-JPG and JSRT datasets, demonstrating superior translation quality through lower Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) scores compared to existing methods (FID: 67.18 vs. 210.4, KID: 0.01604 vs. 0.225). Evaluation on RSNA opacity, MIMIC acute respiratory distress syndrome (ARDS) patient CXRs and JSRT CXRs shows our method enhances segmentation accuracy of lung borders and improves lesion classification, further underscoring its potential in clinical settings (RSNA: mIoU: 76.58% vs. 62.58%, Sensitivity: 85.58% vs. 77.03%; MIMIC ARDS: mIoU: 86.20% vs. 72.07%, Sensitivity: 92.68% vs. 86.85%; JSRT: mIoU: 91.08% vs. 85.6%, Sensitivity: 97.62% vs. 95.04%). Our approach advances CXR imaging analysis, especially in investigating segmentation impacts through image translation techniques. • Unpaired translation removes lung opacities yet keeps key features in X-rays. • Adaptive masks highlight and constrain opacity changes for better interpretability. • Cross-domain alignment reduces artefacts and preserves real diagnostic features. • Experiments show improved image fidelity, segmentation, and lesion classification. Junzhi Ning, Dominic C. Marshall, Yijian Gao, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Matthieu Komorowski, Guang Yang 0006 |
Pattern Recognit. Lett. | 7 |
| 2024 | Fuzzy Attention-Based Border Rendering Network for Lung Organ Segmentation
Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Xiaodan Xing, Zhifan Gao, Guang Yang 0006 |
MICCAI (9) | 1 |
| 2024 | Dynamic Multimodal Information Bottleneck for Multimodality ClassificationabstractEffectively leveraging multimodal data such as various images, laboratory tests and clinical information is becoming increasingly attractive in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on enhancing their performance by leveraging the differences or shared features from various modalities and fusing feature across different modalities. These approaches are generally not optimal for clinical settings, which pose the additional challenges of limited training data, as well as being rife with redundant data or noisy modality channels, leading to subpar performance. To address this gap, we study the robustness of existing methods to data redundancy and noise and propose a generalized dynamic multimodal information bottleneck framework for attaining a robust fused feature representation. Specifically, our information bottleneck module serves to filter out the task-irrelevant information and noises in the fused feature, and we further introduce a sufficiency loss to prevent dropping of task-relevant information, thus explicitly preserving the sufficiency of prediction information in the distilled feature. We validate our model on an in-house and a public COVID19 dataset for mortality prediction as well as two public biomedical datasets for diagnostic tasks. Extensive experiments show that our method surpasses the state-of-the-art and is significantly more robust, being the only method to remain performance when large-scale noisy channels exist. Our code is publicly available at https://github.com/ayanglab/DMIB. Yingying Fang, Shuang Wu 0002, Sheng Zhang 0024, Chaoyan Huang, Tieyong Zeng, Xiaodan Xing, Simon Walsh, Guang Yang 0006 |
WACV | 3 |
| 2024 | Probing perfection: The relentless art of meddling for pulmonary airway segmentation from HRCT via a human-AI collaboration based active learning methodabstractIn the realm of pulmonary tracheal segmentation, the scarcity of annotated data stands as a prevalent pain point in most medical segmentation endeavors. Concurrently, most Deep Learning (DL) methodologies employed in this domain invariably grapple with other dual challenges: the inherent opacity of 'black box' models and the ongoing pursuit of performance enhancement. In response to these intertwined challenges, the core concept of our Human-Computer Interaction (HCI) based learning models (RS_UNet, LC_UNet, UUNet and WD_UNet) hinge on the versatile combination of diverse query strategies and an array of deep learning models. We train four HCI models based on the initial training dataset and sequentially repeat the following steps 1-4: (1) Query Strategy: Our proposed HCI models selects those samples which contribute the most additional representative information when labeled in each iteration of the query strategy (showing the names and sequence numbers of the samples to be annotated). Additionally, in this phase, the model selects the unlabeled samples with the greatest predictive disparity by calculating the Wasserstein Distance, Least Confidence, Entropy Sampling, and Random Sampling. (2) Central line correction: The selected samples in previous stage are then used for domain expert correction of the system-generated tracheal central lines in each training round. (3) Update training dataset: When domain experts are involved in each epoch of the DL model's training iterations, they update the training dataset with greater precision after each epoch, thereby enhancing the trustworthiness of the 'black box' DL model and improving the performance of models. (4) Model training: Proposed HCI model is trained using the updated training dataset and an enhanced version of existing UNet. Experimental results validate the effectiveness of this Human-Computer Interaction-based approaches, demonstrating that our proposed WD-UNet, LC-UNet, UUNet, RS-UNet achieve comparable or even superior performance than the state-of-the-art DL models, such as WD-UNet with only 15 %-35 % of the training data, leading to substantial reductions (65 %-85 % reduction of annotation effort) in physician annotation time. Yang Nan 0002, Sheng Zhang 0024, Federico Felder, Xiaodan Xing, Yingying Fang, Javier Del Ser, Simon Walsh, Guang Yang 0006 |
Artif. Intell. Medicine | 3 |
| 2024 | Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challengeabstract• This paper investigates the capacity of AI models for airway modelling on national datasets with paired clinical metadata. • We evaluated AI models against unharmonised, noisy, and out-of-distribution data, as well as the prognostication for FLD. • We found a new biomarker for mortality prediction, outperforming existing clinical measurements (FVC% and fibrosis scores). • In-depth analysis of AI models on airway modelling and prognosis, highlighting challenges and future research directions. Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway structures remains prohibitively time-consuming. While significant efforts have been made towards enhancing automatic airway modelling, current public-available datasets predominantly concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023′ (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for mortality prediction, a strong airway-derived biomarker (Hazard ratio>1.5, p < 0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers. Yang Nan 0002, Xiaodan Xing, Zeyu Tang 0001, Federico Felder, Sheng Zhang 0024, Roberta Eufrasia Ledda, Xiaoliu Ding, Feng Shi 0001, Tianyang Sun, Zehong Cao, Yun Gu, Pingyu Wang, Wen Tang 0005, Pengxin Yu, Han Kang, Junqiang Chen, Michail Mamalakis, Francesco Prinzi, Gianluca Carlini, Lisa Cuneo, Abhirup Banerjee, Zhaohu Xing, Lei Zhu 0003, Zacharia Mesbah, Dhruv Jain, Tsiry Mayet, Hongyu Yuan, Qing Lyu 0009, Abdul Qayyum 0002, Moona Mazher, Athol Wells, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 6 |
| 2024 | Fuzzy Attention-Based Border Rendering Orthogonal Network for Lung Organ SegmentationabstractAutomatic lung organ segmentation on computerized tomography images is crucial for lung disease diagnosis. However, the unlimited voxel values and class imbalance of lung organs can lead to false-negative/positive and leakage issues in numerous state-of-the-art methods. In addition, some lung organs are easily lost during therecycleddown/up-sample procedure, e.g., bronchioles and arterioles, which can cause severe discontinuity issue. Inspired by these, this article introduces an effective lung organ segmentation method called fuzzy attention-based border rendering feature orthogonal network, which 1) integrates an efficient transformer-like fuzzy-attention module into deep networks to cope with the uncertainty in feature representations; 2) decouples and depicts the lung organ regions as cube-trees by focusing only onrecycle-sampling border vulnerable points, rendering the severely discontinuous, false-negative/positive organ regions with two novel global-local cube-tree fusion and sparse patched feature orthogonal modules; 3) develops a multiscale self-knowledge guidance module to improve model performance and robustness. We have demonstrated the efficacy of proposed method on five challenging datasets of lung organ segmentation, i.e., airway and artery. All experimental results demonstrate that our method can achieve the favorable performance significantly. Sheng Zhang 0024, Yingying Fang, Yang Nan 0002, Weiping Ding 0001, Yew-Soon Ong, Alejandro F. Frangi, Witold Pedrycz, Simon Walsh, Guang Yang 0006 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2021 | OPMP: An Omnidirectional Pyramid Mask Proposal Network for Arbitrary-Shape Scene Text DetectionabstractScene text detection methods have achieved significant progresses. However, stack-omnidirectional text dilemma, under-segmentation of very close text words, and over-segmentation of arbitrary-shape long text lines, are still main challenges. Motivated by these problems, we proposed a two stage method called omnidirectional pyramid mask proposal text detector (OPMP). OPMP removes anchor mechanism that requires heuristic non-maximum suppress processing. Instead, it uses an effective pyramid lengthwise and sidewise residual sequence modeling method to produce arbitrary-shape proposals. To accurately extract the features of text shape, OPMP enhances the backbone layers by a multiple arbitrary-shape fitting mechanism. Finally, a multi-grain text classification module is proposed, which reclassifies each text region robustly. Comprehensive ablation studies demonstrate the effectiveness of each proposed component. In addition, experiments on various benchmarks, including ICDAR2015, MLT, MSRA-TD500, CTW1500, and Total-text, show that our method outperforms previous state-of-the-art methods. Sheng Zhang 0024, Zhongrong Wei, Chunhua Shen |
IEEE Trans. Multim. | 1 |
| 2019 | Omnidirectional Scene Text Detection with Sequential-free Box DiscretizationabstractScene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection methods. However, recent researches reveal that introducing quadrilateral bounding box for scene text detection will bring a label confusion issue which is easily overlooked, and this issue may significantly undermine the detection performance. To address this issue, in this paper, we propose a novel method called Sequential-free Box Discretization (SBD) by discretizing the bounding box into key edges (KE) which can further derive more effective methods to improve detection performance. Experiments showed that the proposed method can outperform state-of-the-art methods in many popular scene text benchmarks, including ICDAR 2015, MLT, and MSRA-TD500. Ablation study also showed that simply integrating the SBD into Mask R-CNN framework, the detection performance can be substantially improved. Furthermore, an experiment on the general object dataset HRSC2016 (multi-oriented ships) showed that our method can outperform recent state-of-the-art methods by a large margin, demonstrating its powerful generalization ability. Sheng Zhang 0024, Lele Xie, Yaqiang Wu, Zhepeng Wang 0002 |
IJCAI | 2 |
| 2019 | Curved scene text detection via transverse and longitudinal sequence connectionabstractCurved text detection is a difficult problem that has not been addressed sufficiently. To highlight the difficulties in reading curved text in a real environment, we constructed a curved text dataset called CTW1500, which includes over 10,000 text annotations in 1500 images, and used it to formulate a polygon-based curved text detector that can detect curved text without using an empirical combination. With the seamless integration of recurrent transverse and longitudinal offset connection, our method explores context information instead of predicting points independently, resulting in smoother and more accurate detection. Our approach is designed as a universal method, meaning it can be trained using rectangular or quadrilateral bounding boxes, requiring no extra effort. Experimental results on the CTW1500 dataset and Total-text demonstrated that our method with only a light backbone can outperform state-of-the-art methods by a large margin. Our method also achieved state-of-the-art performance on the MSRA-TD500 dataset, demonstrating its promising generalization ability. Code, datasets, and label-tool are available at https://github.com/Yuliang-Liu/Curve-Text-Detector. Shuaitao Zhang, Canjie Luo, Sheng Zhang 0024 |
Pattern Recognit. | 5 |
| 2018 | Feature Enhancement Network: A Refined Scene Text DetectorabstractIn this paper, we propose a refined scene text detector with a novel Feature Enhancement Network (FEN)for Region Proposal and Text Detection Refinement. Retrospectively, both region proposal with only 3 x 3 sliding-window feature and text detection refinement with single scale high level feature are insufficient, especially for smaller scene text. Therefore, we design a new FEN network with task-specific, low and high level semantic features fusion to improve the performance of text detection. Besides, since unitary position-sensitive RoI pooling in general object detection is unreasonable for variable text regions, an adaptively weighted position-sensitive RoI pooling layer is devised for further enhancing the detecting accuracy. To tackle the sample-imbalance problem during the refinement stage,we also propose an effective positives mining strategy for efficiently training our network. Experiments on ICDAR2011 and 2013 robust text detection benchmarks demonstrate that our method can achieve state-of-the-art results, outperforming all reported methods in terms of F-measure. Sheng Zhang 0024, Canjie Luo |
AAAI | 1 |
| 2018 | ICPR2018 Contest on Robust Reading for Multi-Type Web ImagesabstractElectronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results. Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013 |
ICPR | 4 |
| 2018 | A New CNN-Based Method for Multi-Directional Car License Plate DetectionabstractThis paper presents a novel convolutional neural network (CNN) -based method for high-accuracy real-time car license plate detection. Many contemporary methods for car license plate detection are reasonably effective under the specific conditions or strong assumptions only. However, they exhibit poor performance when the assessed car license plate images have a degree of rotation, as a result of manual capture by traffic police or deviation of the camera. Therefore, we propose the a CNN-based MD-YOLO framework for multi-directional car license plate detection. Using accurate rotation angle prediction and a fast intersection-over-union evaluation strategy, our proposed method can elegantly manage rotational problems in real-time scenarios. A series of experiments have been carried out to establish that the proposed method outperforms over other existing state-of-the-art methods in terms of better accuracy and lower computational cost. Lele Xie, Tasweer Ahmad, Sheng Zhang 0024 |
IEEE Trans. Intell. Transp. Syst. | 5 |