Sukun Tian

dblp:232/6435 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-8289-5968ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Few-Shot Pulmonary Vessel Segmentation Based on Tubular-Aware Prompt-Tuning
abstract
Segmentation of the pulmonary vessel from computed tomography (CT) images plays a crucial role in the diagnosis and treatment of various lung diseases. Although deep learning-based approaches have shown remarkable progress in recent years, their performance is often hindered by the lack of high-quality annotated datasets, in which the complex anatomy and morphology of pulmonary vessels make manual annotation challenging, time-consuming, and prone to errors. To address this, we propose PV25, the first dataset that features finely paired annotations of both pulmonary vessels and airways. Moreover, we propose TPNet, a novel tubular-aware prompt-tuning framework for pulmonary vessel segmentation under few-shot training with limited annotations. Specifically, based on an advanced and frozen segmentation backbone, TPNet proposes tunable encoding and decoding networks that learn tubular structures as transfer learning priors, bridging the gap between the source and target pulmonary vessel domains. Specifically, TPNet is built in an encoder-decoder manner, including the fixed segmentation backbone, tunable encoding and decoding networks. In encoding stage, the Morphology-Driven Region Growing (MDRG) module is developed to leverage the tubular connectivity of vessels to guide the network in capturing fine-grained features of pulmonary vessels. In decoding stage, the Cross-Correlation Guidance (CCG) module is introduced to integrate multi-scale correlations between airway and vessel structures in a coarse-to-fine manner. Extensive experiments conducted on multiple datasets demonstrate that TPNet achieves state-of-the-art performance in pulmonary vessel segmentation under limited training data. Besides, TPNet shows strong performance in related tasks such as airway segmentation and artery-vein classification, highlighting its robustness and versatility.
Zijian Gao, Lai Jiang 0004, Sukun Tian, Yuchun Sun, Mai Xu, Liyuan Tao
IEEE Trans. Medical Imaging4
2024 RADDA-Net: Residual attention-based dual discriminator adversarial network for surface defect detection
Sukun Tian, Pan Huang 0001, Renkai Huang
Eng. Appl. Artif. Intell.1
2024 LA-ViT: A Network With Transformers Constrained by Learned-Parameter-Free Attention for Interpretable Grading in a New Laryngeal Histopathology Image Dataset
abstract
Grading laryngeal squamous cell carcinoma (LSCC) based on histopathological images is a clinically significant yet challenging task. However, more low-effect background semantic information appeared in the feature maps, feature channels, and class activation maps, which caused a serious impact on the accuracy and interpretability of LSCC grading. While the traditional transformer block makes extensive use of parameter attention, the model overlearns the low-effect background semantic information, resulting in ineffectively reducing the proportion of background semantics. Therefore, we propose an end-to-end network with transformers constrained by learned-parameter-free attention (LA-ViT), which improve the ability to learn high-effect target semantic information and reduce the proportion of background semantics. Firstly, according to generalized linear model and probabilistic, we demonstrate that learned-parameter-free attention (LA) has a stronger ability to learn highly effective target semantic information than parameter attention. Secondly, the first-type LA transformer block of LA-ViT utilizes the feature map position subspace to realize the query. Then, it uses the feature channel subspace to realize the key, and adopts the average convergence to obtain a value. And those construct the LA mechanism. Thus, it reduces the proportion of background semantics in the feature maps and feature channels. Thirdly, the second-type LA transformer block of LA-ViT uses the model probability matrix information and decision level weight information to realize key and query, respectively. And those realize the LA mechanism. So, it reduces the proportion of background semantics in class activation maps. Finally, we build a new complex semantic LSCC pathology image dataset to address the problem, which is less research on LSCC grading models because of lacking clinically meaningful datasets. After extensive experiments, the whole metrics of LA-ViT outperform those of other state-of-the-art methods, and the visualization maps match better with the regions of interest in the pathologists' decision-making. Moreover, the experimental results conducted on a public LSCC pathology image dataset show that LA-ViT has superior generalization performance to that of other state-of-the-art methods.
Pan Huang 0001, Hualiang Xiao, Peng He 0002, Chentao Li 0002, Sukun Tian, Peng Feng 0002, Yuchun Sun, Francesco Mercaldo, Antonella Santone, Harry Qin
IEEE J. Biomed. Health Informatics6
2023 Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven Approach
abstract
Just noticeable difference (JND) refers to the maximum visual change that human eyes cannot perceive, and it has a wide range of applications in multimedia systems. However, most existing JND approaches only focus on a single modality, and rarely consider the complementary effects of multimodal information. In this article, we investigate the JND modeling from an end-to-end homologous multimodal perspective, namely hmJND-Net. Specifically, we explore three important visually sensitive modalities, including saliency, depth, and segmentation. To better utilize homologous multimodal information, we establish an effective fusion method via summation enhancement and subtractive offset, and align homologous multimodal features based on a self-attention driven encoder-decoder paradigm. Extensive experimental results on eight different benchmark datasets validate the superiority of our hmJND-Net over eight representative methods.
Wuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang, Ye Liu 0005, Miaohui Wang
AAAI3
2023 TranSDFNet: Transformer-Based Truncated Signed Distance Fields for the Shape Design of Removable Partial Denture Clasps
abstract
The ever-growing aging population has led to an increasing need for removable partial dentures (RPDs) since they are typically the least expensive treatment options for partial edentulism. However, the digital design of RPDs remains challenging for dental technicians due to the variety of partially edentulous scenarios and complex combinations of denture components. To accelerate the design of RPDs, we propose a U-shape network incorporated with Transformer blocks to automatically generate RPD clasps, one of the most frequently used RPD components. Unlike existing dental restoration design algorithms, we introduce the voxel-based truncated signed distance field (TSDF) as an intermediate representation, which reduces the sensitivity of the network to resolution and contributes to more smooth reconstruction. Besides, a selective insertion scheme is proposed for solving the memory issue caused by Transformer blocks and enables the algorithm to work well in scenarios with insufficient data. We further design two weighted loss functions to filter out the noisy signals generated from the zero-gradient areas in TSDF. Ablation and comparison studies demonstrate that our algorithm outperforms state-of-the-art reconstruction methods by a large margin and can serve as an intelligent auxiliary in denture design.
Xinze Shen, Changdong Zhang, Xiuyi Jia, Sukun Tian, Yuchun Sun, Wenhe Liao
IEEE J. Biomed. Health Informatics6
2023 A ViT-AMC Network With Adaptive Model Fusion and Multiobjective Optimization for Interpretable Laryngeal Tumor Grading From Histopathological Images
abstract
The tumor grading of laryngeal cancer pathological images needs to be accurate and interpretable. The deep learning model based on the attention mechanism-integrated convolution (AMC) block has good inductive bias capability but poor interpretability, whereas the deep learning model based on the vision transformer (ViT) block has good interpretability but weak inductive bias ability. Therefore, we propose an end-to-end ViT-AMC network (ViT-AMCNet) with adaptive model fusion and multiobjective optimization that integrates and fuses the ViT and AMC blocks. However, existing model fusion methods often have negative fusion: 1). There is no guarantee that the ViT and AMC blocks will simultaneously have good feature representation capability. 2). The difference in feature representations learning between the ViT and AMC blocks is not obvious, so there is much redundant information in the two feature representations. Accordingly, we first prove the feasibility of fusing the ViT and AMC blocks based on Hoeffding's inequality. Then, we propose a multiobjective optimization method to solve the problem that ViT and AMC blocks cannot simultaneously have good feature representation. Finally, an adaptive model fusion method integrating the metrics block and the fusion block is proposed to increase the differences between feature representations and improve the deredundancy capability. Our methods improve the fusion ability of ViT-AMCNet, and experimental results demonstrate that ViT-AMCNet significantly outperforms state-of-the-art methods. Importantly, the visualized interpretive maps are closer to the region of interest of concern by pathologists, and the generalization ability is also excellent. Our code is publicly available at https://github.com/Baron-Huang/ViT-AMCNet.
Pan Huang 0001, Peng He 0002, Sukun Tian, Peng Feng 0002, Hualiang Xiao, Francesco Mercaldo, Antonella Santone, Harry Qin
IEEE Trans. Medical Imaging3
2022 DCPR-GAN: Dental Crown Prosthesis Restoration Using Two-Stage Generative Adversarial Networks
abstract
Restoring the correct masticatory function of broken teeth is the basis of dental crown prosthesis rehabilitation. However, it is a challenging task primarily due to the complex and personalized morphology of the occlusal surface. In this article, we address this problem by designing a new two-stage generative adversarial network (GAN) to reconstruct a dental crown surface in the data-driven perspective. Specifically, in the first stage, a conditional GAN (CGAN) is designed to learn the inherent relationship between the defective tooth and the target crown, which can solve the problem of the occlusal relationship restoration. In the second stage, an improved CGAN is further devised by considering an occlusal groove parsing network (GroNet) and an occlusal fingerprint constraint to enforce the generator to enrich the functional characteristics of the occlusal surface. Experimental results demonstrate that the proposed framework significantly outperforms the state-of-the-art deep learning methods in functional occlusal surface reconstruction using a real-world patient database. Moreover, the standard deviation (SD) and root mean square (RMS) between the generated occlusal surface and the target crown calculated by our method are both less than 0.161 mm. Importantly, the designed dental crown have enough anatomical morphology and higher clinical applicability.
Sukun Tian, Miaohui Wang, Luca Fiorenza, Yuchun Sun, Yangmin Li 0001
IEEE J. Biomed. Health Informatics1
2021 Efficient Computer-Aided Design of Dental Inlay Restoration: A Deep Adversarial Framework
abstract
Restoring the normal masticatory function of broken teeth is a challenging task primarily due to the defect location and size of a patient's teeth. In recent years, although some representative image-to-image transformation methods (e.g. Pix2Pix) can be potentially applicable to restore the missing crown surface, most of them fail to generate dental inlay surface with realistic crown details (e.g. occlusal groove) that are critical to the restoration of defective teeth with varying shapes. In this article, we design a computer-aided Deep Adversarial-driven dental Inlay reStoration (DAIS) framework to automatically reconstruct a realistic surface for a defective tooth. Specifically, DAIS consists of a Wasserstein generative adversarial network (WGAN) with a specially designed loss measurement, and a new local-global discriminator mechanism. The local discriminator focuses on missing regions to ensure the local consistency of a generated occlusal surface, while the global discriminator aims at defective teeth and adjacent teeth to assess if it is coherent as a whole. Experimental results demonstrate that DAIS is highly efficient to deal with a large area of missing teeth in arbitrary shapes and generate realistic occlusal surface completion. Moreover, the designed watertight inlay prostheses have enough anatomical morphology, thus providing higher clinical applicability compared with more state-of-the-art methods.
Sukun Tian, Miaohui Wang, Fulai Yuan, Yuchun Sun, Wuyuan Xie, Harry Qin
IEEE Trans. Medical Imaging1