VLDB 2026 Research / reviewers in the wild / expert
Yang Zhou 0036
dblp:07/4580-36
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-2848-7642ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniPET: A universal network for high-quality PET image denoising across varied dose reduction factors
Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001 |
Medical Image Anal. | 2 |
| 2026 | VQPET: Leveraging Vector-Quantized Codebook Prior for PET Image SynthesisabstractPositron emission tomography (PET) image synthesis is a highly ill-posed problem that requires auxiliary priors to 1) alleviate the loss of high-quality (HQ) information in low-quality (LQ) inputs, and 2) impose additional constraints to reduce mapping uncertainty. However, existing auxiliary priors in PET image synthesis often provide inadequate guidance due to inaccurate prior information or limited prior expressiveness. To overcome the aforementioned limitations, the vector-quantized (VQ) codebook prior is employed as a promising solution. By learning discrete latent feature representations of HQ images through deep models, the VQ codebook prior encompasses accurate HQ information and possesses great expressiveness. Building upon this, we propose a novel two-stage framework, VQPET, that introduces the VQ codebook prior for PET image synthesis. In the first stage, it pretrains a VQGAN on an additional large-scale HQ PET dataset, encoding intrinsic HQ features as code items in the VQ codebook. The VQ codebook prior is thus derived from the high-level features obtained from the pretrained VQGAN and serves as an additional constraint for downstream synthesis. In the second stage, it develops a codebook-prior-guided network (CPGNet) that effectively exploits the VQ codebook prior to produce realistic outputs. Specifically, CPGNet progressively incorporates the VQ codebook prior at multiple decoding levels, providing reliable guidance for HQ synthesis. Compared to previous works, VQPET innovatively leverages additional large-scale HQ datasets to transfer pretrained prior knowledge for enhanced synthesis and functions as a general framework applicable to any encoder-decoder network. Extensive experiments demonstrate the substantial effect and robust generalizability of VQPET. Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for PathologyabstractMultimodal large language models (MLLMs) have demonstrated strong performance in patch-level pathological image analysis; however, they often lack the holistic perceptual capability necessary for comprehensive Whole Slide Image (WSI) interpretation. Recent approaches have explored constructing slide-level MLLMs using VQA datasets that are entirely generated from pathology reports by large language models (LLMs). However, these datasets suffer from critical limitations: hallucinated content, information leakage in question stems, clinically irrelevant or visual independent questions, and the omission of essential diagnostic features-issues that undermine both data quality and clinical validity. In this paper, we introduce a clinical diagnosis template-based pipeline to collect pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA-BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS-Align, a dataset of 80k slide-description pairs from 804 WSIs for vision-language alignment training, and CTISBench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question-answer pairs. CTIS-Bench emphasizes clinically grounded, closed-ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non-visual reasoning, and require genuine slide understanding. We further propose CTIS-QA, a Slide-level Question Answering model, featuring a dual-stream architecture that mimics pathologists' diagnostic approach. One stream captures global slidelevel context via clustering-based feature aggregation, while the other focuses on salient local regions through attention-guided patch perception module. Extensive experiments on WSI-VQA, CTIS-Bench, and slide-level diagnostic tasks show that CTIS-QA consistently outperforms existing state-of-the-art models across multiple metrics. We will fully release both CTIS-Bench and CTIS-QA as open-source resources. Ziniu Qian, Yang Zhou 0036, Bingzheng Wei, Yan Xu 0001 |
BIBM | 4 |
| 2025 | Visual Textualization for Image Prompted Object DetectionabstractWe propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are difficult to describe textually and nearly absent from their pre-training data, while preserving their pre-trained object-text alignment. Specifically, VisTex-OVLM leverages multi-scale textualizing blocks and a multi-stage fusion strategy to integrate visual information from visual exemplars, generating textualized visual tokens that effectively guide OVLMs alongside text prompts. Unlike previous methods, our method maintains the original architecture of OVLM, maintaining its generalization capabilities while enhancing performance in few-shot settings. VisTex-OVLM demonstrates superior performance across open-set datasets which have minimal overlap with OVLM's pre-training data and achieves state-of-the-art results on few-shot benchmarks PASCAL VOC and MSCOCO. The code will be released at https://github.com/WitGotFlg/VisTex-OVLM. Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Yan Xu 0001 |
ICCV | 2 |
| 2025 | AttriPrompter: Auto-Prompting With Attribute Semantics for Zero-Shot Nuclei Detection via Visual-Language Pre-Trained ModelsabstractLarge-scale visual-language pre-trained models (VLPMs) have demonstrated exceptional performance in downstream object detection through text prompts for natural scenes. However, their application to zero-shot nuclei detection on histopathology images remains relatively unexplored, mainly due to the significant gap between the characteristics of medical images and the web-originated text-image pairs used for pre-training. This paper aims to investigate the potential of the object-level VLPM, Grounded Language-Image Pre-training (GLIP), for zero-shot nuclei detection. Specifically, we propose an innovative auto-prompting pipeline, named AttriPrompter, comprising attribute generation, attribute augmentation, and relevance sorting, to avoid subjective manual prompt design. AttriPrompter utilizes VLPMs' text-to-image alignment to create semantically rich text prompts, which are then fed into GLIP for initial zero-shot nuclei detection. Additionally, we propose a self-trained knowledge distillation framework, where GLIP serves as the teacher with its initial predictions used as pseudo labels, to address the challenges posed by high nuclei density, including missed detections, false positives, and overlapping instances. Our method exhibits remarkable performance in label-free nuclei detection, outperforming all existing unsupervised methods and demonstrating excellent generality. Notably, this work highlights the astonishing potential of VLPMs pre-trained on natural image-text pairs for downstream tasks in the medical field as well. Code will be released at github.com/AttriPrompter. Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yan Xu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Tuning Stable Rank Shrinkage: Aiming at the Overlooked Structural Risk in Fine-tuningabstractExisting finetuning methods for computer vision tasks primarily focus on re-weighting the knowledge learned from the source domain during pre-training. They aim to retain beneficial knowledge for the target domain while suppressing unfavorable knowledge. During the pre-training and fine-tuning stages, there is a notable disparity in the data scale. Consequently, it is theoretically necessary to employ a model with reduced complexity to mitigate the potential structural risk. However, our empirical investigation in this paper reveals that models finetuned using existing methods still manifest a high level of model complexity inherited from the pre-training stage, leading to a suboptimal stability and generalization ability. This phenomenon indicates an issue that has been overlooked in fine-tuning: Structural Risk Minimization. To address this issue caused by data scale disparity during the fine-tuning stage, we propose a simple yet effective approach called Tuning Stable Rank Shrinkage (TSRS). TSRS mitigates the structural risk during the fine-tuning stage by constraining the noise sensitivity of the target model based on stable rank theories. Through extensive experiments, we demonstrate that incorporating TSRS into fine-tuning methods leads to improved generalization ability on various tasks, regardless of whether the neural networks are based on convolution or transformer architectures. Additionally, empirical analysis reveals that TSRS enhances the robustness, convexity, and smoothness of the loss landscapes in fine-tuned models. Code is available at https://github.com/WitGotFlg/TSRS. Sicong Shen, Yang Zhou 0036, Bingzheng Wei, Eric I-Chao Chang, Yan Xu 0001 |
CVPR | 2 |
| 2024 | SDPT: Synchronous Dual Prompt Tuning for Fusion-Based Visual-Language Pre-trained Models
Yang Zhou 0036, Yongjian Wu 0002, Jiya Saiyin, Bingzheng Wei, Maode Lai, Eric Chang, Yan Xu 0001 |
ECCV (49) | 1 |
| 2024 | Region Attention Transformer for Medical Image Restoration
Zhiwen Yang 0001, Ziniu Qian, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001 |
MICCAI (7) | 4 |
| 2023 | Zero-Shot Nuclei Detection via Visual-Language Pre-trained Models
Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yubo Fan, Yan Xu 0001 |
MICCAI (6) | 2 |
| 2023 | DRMC: A Generalist Model with Dynamic Routing for Multi-center PET Image Synthesis
Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yubo Fan, Yan Xu 0001 |
MICCAI (3) | 2 |
| 2023 | Cyclic Learning: Bridging Image-Level Labels and Nuclei Instance SegmentationabstractNuclei instance segmentation on histopathology images is of great clinical value for disease analysis. Generally, fully-supervised algorithms for this task require pixel-wise manual annotations, which is especially time-consuming and laborious for the high nuclei density. To alleviate the annotation burden, we seek to solve the problem through image-level weakly supervised learning, which is underexplored for nuclei instance segmentation. Compared with most existing methods using other weak annotations (scribble, point, etc.) for nuclei instance segmentation, our method is more labor-saving. The obstacle to using image-level annotations in nuclei instance segmentation is the lack of adequate location information, leading to severe nuclei omission or overlaps. In this paper, we propose a novel image-level weakly supervised method, called cyclic learning, to solve this problem. Cyclic learning comprises a front-end classification task and a back-end semi-supervised instance segmentation task to benefit from multi-task learning (MTL). We utilize a deep learning classifier with interpretability as the front-end to convert image-level labels to sets of high-confidence pseudo masks and establish a semi-supervised architecture as the back-end to conduct nuclei instance segmentation under the supervision of these pseudo masks. Most importantly, cyclic learning is designed to circularly share knowledge between the front-end classifier and the back-end semi-supervised part, which allows the whole system to fully extract the underlying information from image-level labels and converge to a better optimum. Experiments on three datasets demonstrate the good generality of our method, which outperforms other image-level weakly supervised methods for nuclei instance segmentation, and achieves comparable performance to fully-supervised methods. Yang Zhou 0036, Yongjian Wu 0002, Zihua Wang, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yubo Fan, Yan Xu 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2022 | 3D Segmentation Guided Style-Based Generative Adversarial Networks for PET SynthesisabstractPotential radioactive hazards in full-dose positron emission tomography (PET) imaging remain a concern, whereas the quality of low-dose images is never desirable for clinical use. So it is of great interest to translate low-dose PET images into full-dose. Previous studies based on deep learning methods usually directly extract hierarchical features for reconstruction. We notice that the importance of each feature is different and they should be weighted dissimilarly so that tiny information can be captured by the neural network. Furthermore, the synthesis on some regions of interest is important in some applications. Here we propose a novel segmentation guided style-based generative adversarial network (SGSGAN) for PET synthesis. (1) We put forward a style-based generator employing style modulation, which specifically controls the hierarchical features in the translation process, to generate images with more realistic textures. (2) We adopt a task-driven strategy that couples a segmentation task with a generative adversarial network (GAN) framework to improve the translation performance. Extensive experiments show the superiority of our overall framework in PET synthesis, especially on those regions of interest. Yang Zhou 0036, Zhiwen Yang 0001, Hui Zhang 0099, Eric I-Chao Chang, Yubo Fan, Yan Xu 0001 |
IEEE Trans. Medical Imaging | 1 |