Dazhou Guo

dblp:160/2794 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-5591-6126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Preoperative Prediction of Esophageal Cancer Survival in CT via Tumor and Lymph Node Context and Geometry Modeling
abstract
Esophageal cancer is one of the most lethal cancers, with 5-year survival rate of only 20%. Patient outcomes can vary significantly even though they are at the same cancer stage and receive similar treatments. Accurate prognostic prediction for esophageal cancer patients is highly desired to receive personalized precise treatment. Nevertheless, there are very few automated methods yet to fully exploit the preoperative contrast-enhanced computed tomography (CE-CT) imaging for assessing esophageal cancer prognosis. In addition to image patterns, important prognostic factors should encompass tumor size and location, as well as lymph nodes (LNs) involvement, including features such as LN number, size, spatial distribution, and their proximity to tumor. Considering these complexities, we propose a novel Tumor and LN Context-Geometry network for the preoperative prediction of esophageal cancer survival in CE-CT images. Specifically, we 1) focus on learning survival patterns of CT texture via co-attention context modeling at most informative regions, i.e., automatically segmented tumor, LNs and LN-stations; and 2) integrate tumor and LN anatomical and spatial associations into neural geometry modeling for a comprehensive learning of metastatic involvement and tumor invasion to adjacent structures. Empirical studies show our presented framework can improve overall survival prediction performances compared with existing state-of-the-art survival analysis methods, and evidently suggest that incorporating these findings into the existing esophageal cancer staging system would add its clinical values.
Yirui Wang 0002, Haoshen Li, Jiawen Yao, Lianzhen Zhong, Dazhou Guo, Ke Yan 0006, David S. Doermann, Le Lu 0001, Feiran Jiao, Tsung-Ying Ho, Ling Zhang 0002, Abudili Abuduxuku, Xianghua Ye, Dakai Jin
IEEE Trans. Medical Imaging7
2025 Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT Scans
abstract
Segment anything model (SAM) demonstrates strong generalization ability on natural image segmentation. However, its direct adaptation in medical image segmentation tasks shows significant performance drops. It also requires an excessive number of prompt points to obtain a reasonable accuracy. Although quite a few studies explore adapting SAM into medical image volumes, the efficiency of 2D adaptation methods is unsatisfactory and 3D adaptation methods are only capable of segmenting specific organs/tumors. In this work, we propose a comprehensive and scalable 3D SAM model for whole-body CT segmentation, named CT-SAM3D. Instead of adapting SAM, we propose a 3D promptable segmentation model using a (nearly) fully labeled CT dataset. To train CT-SAM3D effectively, ensuring the model's accurate responses to higher-dimensional spatial prompts is crucial, and 3D patch-wise training is required due to GPU memory constraints. Therefore, we propose two key technical developments: 1) a progressively and spatially aligned prompt encoding method to effectively encode click prompts in local 3D space; and 2) a cross-patch prompt scheme to capture more 3D spatial context, which is beneficial for reducing the editing workloads when interactively prompting on large organs. CT-SAM3D is trained using a curated dataset of 1204 CT scans containing 107 whole-body anatomies and extensively validated using five datasets, achieving significantly better results against all previous SAM-derived models.
Heng Guo 0008, Tony C. W. Mok, Dazhou Guo, Ke Yan 0006, Le Lu 0001, Dakai Jin, Minfeng Xu
AAAI5
2025 Metastatic Lymph Node Station Classification in Esophageal Cancer via Prior-Guided Supervision and Station-Aware Mixture-of-Experts
Haoshen Li, Yirui Wang 0002, Qinji Yu, Ke Yan 0006, Dazhou Guo, Le Lu 0001, Bin Dong 0001, Li Zhang 0047, Xianghua Ye, Dakai Jin
MICCAI (13)6
2024 Effective Lymph Nodes Detection in CT Scans Using Location Debiased Query Selection and Contrastive Query Representation in Transformer
Qinji Yu, Yirui Wang 0002, Ke Yan 0006, Haoshen Li, Dazhou Guo, Li Zhang 0047, Na Shen, Le Lu 0001, Xianghua Ye, Dakai Jin
ECCV (42)5
2024 Semi-supervised Lymph Node Metastasis Classification with Pathology-Guided Label Sharpening and Two-Streamed Multi-scale Fusion
Haoshen Li, Yirui Wang 0002, Dazhou Guo, Qinji Yu, Ke Yan 0006, Le Lu 0001, Xianghua Ye, Li Zhang 0047, Dakai Jin
MICCAI (11)4
2024 Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation
Vince Zhu, Zhanghexuan Ji, Dazhou Guo, Puyang Wang, Yingda Xia, Le Lu 0001, Xianghua Ye, Wei Zhu 0015, Dakai Jin
MICCAI (8)3
2024 LViT: Language Meets Vision Transformer in Medical Image Segmentation
abstract
Deep learning has been widely used in medical image segmentation and other aspects. However, the performance of existing medical image segmentation models has been limited by the challenge of obtaining sufficient high-quality labeled data due to the prohibitive data annotation cost. To alleviate this limitation, we propose a new text-augmented medical image segmentation model LViT (Language meets Vision Transformer). In our LViT model, medical text annotation is incorporated to compensate for the quality deficiency in image data. In addition, the text information can guide to generate pseudo labels of improved quality in the semi-supervised learning. We also propose an Exponential Pseudo label Iteration mechanism (EPI) to help the Pixel-Level Attention Module (PLAM) preserve local image features in semi-supervised LViT setting. In our model, LV (Language-Vision) loss is designed to supervise the training of unlabeled images using text information directly. For evaluation, we construct three multimodal medical segmentation datasets (image + text) containing X-rays and CT images. Experimental results show that our proposed LViT has superior segmentation performance in both fully-supervised and semi-supervised setting. The code and datasets are available at https://github.com/HUANGLIZI/LViT.
Qingde Li, Puyang Wang, Dazhou Guo, Le Lu 0001, Dakai Jin, You Zhang 0003, Qingqi Hong
IEEE Trans. Medical Imaging5
2024 Accurate Airway Tree Segmentation in CT Scans via Anatomy-Aware Multi-Class Segmentation and Topology-Guided Iterative Learning
abstract
Intrathoracic airway segmentation in computed tomography is a prerequisite for various respiratory disease analyses such as chronic obstructive pulmonary disease, asthma and lung cancer. Due to the low imaging contrast and noises execrated at peripheral branches, the topological-complexity and the intra-class imbalance of airway tree, it remains challenging for deep learning-based methods to segment the complete airway tree (on extracting deeper branches). Unlike other organs with simpler shapes or topology, the airway's complex tree structure imposes an unbearable burden to generate the "ground truth" label (up to 7 or 3 hours of manual or semi-automatic annotation per case). Most of the existing airway datasets are incompletely labeled/annotated, thus limiting the completeness of computer-segmented airway. In this paper, we propose a new anatomy-aware multi-class airway segmentation method enhanced by topology-guided iterative self-learning. Based on the natural airway anatomy, we formulate a simple yet highly effective anatomy-aware multi-class segmentation task to intuitively handle the severe intra-class imbalance of the airway. To solve the incomplete labeling issue, we propose a tailored iterative self-learning scheme to segment toward the complete airway tree. For generating pseudo-labels to achieve higher sensitivity (while retaining similar specificity), we introduce a novel breakage attention map and design a topology-guided pseudo-label refinement method by iteratively connecting breaking branches commonly existed from initial pseudo-labels. Extensive experiments have been conducted on four datasets including two public challenges. The proposed method achieves the top performance in both EXACT'09 challenge using average score and ATM'22 challenge on weighted average score. In a public BAS dataset and a private lung cancer dataset, our method significantly improves previous leading approaches by extracting at least (absolute) 6.1% more detected tree length and 5.2% more tree branches, while maintaining comparable precision.
Puyang Wang, Dazhou Guo, Haogang Yu, Jia Ge, Yun Gu, Le Lu 0001, Xianghua Ye, Dakai Jin
IEEE Trans. Medical Imaging2
2023 Continual Segment: Towards a Single, Unified and Non-forgetting Continual Segmentation Model of 143 Whole-body Organs in CT Scans
abstract
Deep learning empowers the mainstream medical image segmentation methods. Nevertheless, current deep segmentation approaches are not capable of efficiently and effectively adapting and updating the trained models when new segmentation classes are incrementally added. In the real clinical environment, it can be preferred that segmentation models could be dynamically extended to segment new organs/tumors without the (re-)access to previous training datasets due to obstacles of patient privacy and data storage. This process can be viewed as a continual semantic segmentation (CSS) problem, being understudied for multi-organ segmentation. In this work, we propose a new architectural CSS learning framework to learn a single deep segmentation model for segmenting a total of 143 whole-body organs. Using the encoder/decoder network structure, we demonstrate that a continually trained then frozen encoder coupled with incrementally-added decoders can extract sufficiently representative image features for new classes to be subsequently and validly segmented, while avoiding the catastrophic forgetting in CSS. To maintain a single network model complexity, each decoder is progressively pruned using neural architecture search and teacher-student based knowledge distillation. Finally, we propose a body-part and anomaly-aware output merging module to combine organ predictions originating from different decoders and incorporate both healthy and pathological organs appearing in different datasets. Trained and validated on 3D CT scans of 2500+ patients from four datasets, our single network can segment a total of 143 whole-body organs with very high accuracy, closely reaching the upper bound performance level by training four separate segmentation models (i.e., one model per dataset/task).
Zhanghexuan Ji, Dazhou Guo, Puyang Wang, Ke Yan 0006, Le Lu 0001, Minfeng Xu, Jia Ge, Mingchen Gao, Xianghua Ye, Dakai Jin
ICCV2
2023 Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu
Medical Image Anal.16
2022 Thoracic Lymph Node Segmentation in CT Imaging via Lymph Node Station Stratification and Size Encoding
Dazhou Guo, Jia Ge, Ke Yan 0006, Puyang Wang, Zhuotun Zhu, Xian-Sheng Hua 0001, Le Lu 0001, Tsung-Ying Ho, Xianghua Ye, Dakai Jin
MICCAI (5)1
2022 SAM: Self-Supervised Learning of Pixel-Wise Anatomical Embeddings in Radiological Images
abstract
Radiological images such as computed tomography (CT) and X-rays render anatomy with intrinsic structures. Being able to reliably locate the same anatomical structure across varying images is a fundamental task in medical image analysis. In principle it is possible to use landmark detection or semantic segmentation for this task, but to work well these require large numbers of labeled data for each anatomical structure and sub-structure of interest. A more universal approach would learn the intrinsic structure from unlabeled images. We introduce such an approach, called Self-supervised Anatomical eMbedding (SAM). SAM generates semantic embeddings for each image pixel that describes its anatomical location or body part. To produce such embeddings, we propose a pixel-level contrastive learning framework. A coarse-to-fine strategy ensures both global and local anatomical information are encoded. Negative sample selection strategies are designed to enhance the embedding's discriminability. Using SAM, one can label any point of interest on a template image and then locate the same body part in other images by simple nearest neighbor searching. We demonstrate the effectiveness of SAM in multiple tasks with 2D and 3D image modalities. On a chest CT dataset with 19 landmarks, SAM outperforms widely-used registration algorithms while only taking 0.23 seconds for inference. On two X-ray datasets, SAM, with only one labeled template image, surpasses supervised methods trained on 50 labeled images. We also apply SAM on whole-body follow-up lesion matching in CT and obtain an accuracy of 91%. SAM can also be applied for improving image registration and initializing CNN weights.
Ke Yan 0006, Jinzheng Cai, Dakai Jin, Shun Miao, Dazhou Guo, Adam P. Harrison, Youbao Tang, Jing Xiao 0006, Jingjing Lu, Le Lu 0001
IEEE Trans. Medical Imaging5
2021 DeepStationing: Thoracic Lymph Node Station Parsing in CT Scans Using Anatomical Context Encoding and Key Organ Auto-Search
Dazhou Guo, Xianghua Ye, Jia Ge, Xing Di, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Zhongjie Lu, Senxiang Yan, Dakai Jin
MICCAI (5)1
2021 SAME: Deformable Image Registration Based on Self-supervised Anatomical Embeddings
Fengze Liu, Ke Yan 0006, Adam P. Harrison, Dazhou Guo, Le Lu 0001, Alan L. Yuille, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Xianghua Ye, Dakai Jin
MICCAI (4)4
2021 DeepTarget: Gross tumor and clinical target volume segmentation in esophageal cancer radiotherapy
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001
Medical Image Anal.2
2020 Organ at Risk Segmentation for Head and Neck Cancer Using Stratified Learning and Neural Architecture Search
abstract
OAR segmentation is a critical step in radiotherapy of head and neck (H&N) cancer, where inconsistencies across radiation oncologists and prohibitive labor costs motivate automated approaches. However, leading methods using standard fully convolutional network workflows that are challenged when the number of OARs becomes large, e.g. > 40. For such scenarios, insights can be gained from the stratification approaches seen in manual clinical OAR delineation. This is the goal of our work, where we introduce stratified organ at risk segmentation (SOARS), an approach that stratifies OARs into anchor, mid-level, and small & hard (S&H) categories. SOARS stratifies across two dimensions. The first dimension is that distinct processing pipelines are used for each OAR category. In particular, inspired by clinical practices, anchor OARs are used to guide the mid-level and S&H categories. The second dimension is that distinct network architectures are used to manage the significant contrast, size, and anatomy variations between different OARs. We use differentiable neural architecture search (NAS), allowing the network to choose among 2D, 3D or Pseudo-3D convolutions. Extensive 4-fold cross-validation on 142 H&N cancer patients with 42 manually labeled OARs, the most comprehensive OAR dataset to date, demonstrates that both pipeline- and NAS-stratification significantly improves quantitative performance over the state-of-the-art (from 69.52% to 73.68% in absolute Dice scores). Thus, SOARS provides a powerful and principled means to manage the highly complex segmentation space of OARs.
Dazhou Guo, Dakai Jin, Zhuotun Zhu, Tsung-Ying Ho, Adam P. Harrison, Chun-Hung Chao, Jing Xiao 0006, Le Lu 0001
CVPR1
2020 Lymph Node Gross Tumor Volume Detection in Oncology Imaging via Relationship Learning Using Graph Neural Network
Chun-Hung Chao, Zhuotun Zhu, Dazhou Guo, Ke Yan 0006, Tsung-Ying Ho, Jinzheng Cai, Adam P. Harrison, Xianghua Ye, Jing Xiao 0006, Alan L. Yuille, Min Sun 0001, Le Lu 0001, Dakai Jin
MICCAI (7)3
2020 Lymph Node Gross Tumor Volume Detection and Segmentation via Distance-Based Gating Using 3D CT/PET Imaging in Radiotherapy
Zhuotun Zhu, Dakai Jin, Ke Yan 0006, Tsung-Ying Ho, Xianghua Ye, Dazhou Guo, Chun-Hung Chao, Jing Xiao 0006, Alan L. Yuille, Le Lu 0001
MICCAI (7)6
2020 Weakly supervised easy-to-hard learning for object detection in image sequences
Hongkai Yu, Dazhou Guo, Zhipeng Yan, Lan Fu, Jeff P. Simmons, Craig Przybyla, Song Wang 0002
Neurocomputing2
2020 Degraded Image Semantic Segmentation With Dense-Gram Networks
abstract
Degraded image semantic segmentation is of great importance in autonomous driving, highway navigation systems, and many other safety-related applications and it was not systematically studied before. In general, image degradations increase the difficulty of semantic segmentation, usually leading to decreased semantic segmentation accuracy. Therefore, performance on the underlying clean images can be treated as an upper bound of degraded image semantic segmentation. While the use of supervised deep learning has substantially improved the state of the art of semantic image segmentation, the gap between the feature distribution learned using the clean images and the feature distribution learned using the degraded images poses a major obstacle in improving the degraded image semantic segmentation performance. The conventional strategies for reducing the gap include: 1) Adding image-restoration based pre-processing modules; 2) Using both clean and the degraded images for training; 3) Fine-tuning the network pre-trained on the clean image. In this paper, we propose a novel Dense-Gram Network to more effectively reduce the gap than the conventional strategies and segment degraded images. Extensive experiments demonstrate that the proposed Dense-Gram Network yields stateof-the-art semantic segmentation performance on degraded images synthesized using PASCAL VOC 2012, SUNRGBD, CamVid, and CityScapes datasets.
Dazhou Guo, Yanting Pei, Hongkai Yu, Song Wang 0002
IEEE Trans. Image Process.1
2019 Accurate Esophageal Gross Tumor Volume Segmentation in PET/CT Using Two-Stream Chained 3D Deep Network Fusion
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001
MICCAI (2)2
2019 Deep Esophageal Clinical Target Volume Delineation Using Encoded 3D Spatial Context of Tumors, Lymph Nodes, and Organs At Risk
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001
MICCAI (6)2
2019 An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002
Neurocomputing4
2019 Small Object Sensitive Segmentation of Urban Street Scene With Spatial Adjacency Between Object Classes
abstract
Recent advancements in deep learning have shown exciting promise in the urban street scene segmentation. However, many objects, such as poles and sign symbols, are relatively small and they usually cannot be accurately segmented since the larger objects usually contribute more to the segmentation loss. In this paper, we propose a new boundary-based metric that measures the level of spatial adjacency between each pair of object classes and find that this metric is robust against object size induced biases. We develop a new method to enforce this metric into the segmentation loss. We propose a network, which starts with a segmentation network, followed by a new encoder to compute the proposed boundary-based metric, and then trains this network in an end-to-end fashion. In deployment, we only use the trained segmentation network, without the encoder, to segment new unseen images. Experimentally, we evaluate the proposed method using CamVid and CityScapes datasets and achieve a favorable overall performance improvement and a substantial improvement in segmenting small objects.
Dazhou Guo, Ligeng Zhu, Hongkai Yu, Song Wang 0002
IEEE Trans. Image Process.1
2017 Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable Cameras
abstract
In this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by optical flow. However, optical flow is view-variant - the same person's optical flow in different videos can be very different due to view angle change. In this paper, we attempt to utilize 3D human-skeleton sequences to learn a model that can extract view-invariant motion features from optical flows in different views. For this purpose, we use 3D Mocap database to build a synthetic optical flow dataset and train a Triplet Network (TN) consisting of three sub-networks: two for optical flow sequences from different views and one for the underlying 3D Mocap skeleton sequence. Finally, sub-networks for optical flows are used to extract view-invariant features for CVPI. Experimental results show that, using only the motion information, the proposed method can achieve comparable performance with the state-of-the-art methods. Further combination of the proposed method with an appearance-based method achieves new state-of-the-art performance.
Xiaochuan Fan, Yuewei Lin, Hao Guo 0002, Hongkai Yu, Dazhou Guo, Song Wang 0002
ICCV6
2017 Lesion detection using T1-weighted MRI: A new approach based on functional cortical ROIs
abstract
Accurate and precise detection of brain lesions on MR images is important for relating lesion locations to impaired behaviors. In this paper, we propose a method to detect lesion voxels on each functional cortical ROI (Region of Interest) independently using only T1-weighted MR images (T1-MRI). In contrast to existing automatic lesion detection methods, which typically detect lesion voxels on the whole MR image or on gray matter (GM)/white matter (WM), we show that the proposed functional cortical ROI based method can lead to better lesion-detection performance. We evaluate the proposed method using an in-house dataset with 60 chronic stroke patients. Using leave-one-subject-out cross validation, the proposed method can achieve an average Dice coefficient of 0.74 ± 0.11 and outperform three state-of-the-art methods by more than 0.05.
Dazhou Guo, Song Wang 0002
ICIP1