EDBT 2026 Demo / reviewers in the wild / expert
Kuang Gong
dblp:207/0305
· DBLP profile ↗
21ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-2669-2610ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LDM-Morph: Latent diffusion model guided deformable image registrationabstract• LDM-Morph improves medical image registration using a latent diffusion model. • A dual-stream encoder is proposed to enhance feature learning for image alignment. • LGCA module is designed to boost feature interaction and improve precision. • Novel hierarchical similarity metric improves accuracy and topology preservation. Deformable image registration plays an essential role in various medical image tasks. Existing deep learning-based deformable registration frameworks primarily utilize convolutional neural networks (CNNs) or Transformers to learn features to predict the deformations. However, the lack of semantic information in the learned features limits the registration performance. Furthermore, the similarity metric of the loss function is often evaluated only in the pixel space, which ignores the matching of high-level anatomical features and can lead to deformation folding. To address these issues, in this work, we proposed LDM-Morph, an unsupervised deformable registration algorithm for medical image registration. LDM-Morph integrated features extracted from the latent diffusion model (LDM) to enrich the semantic information. Additionally, a latent and global feature-based cross-attention module (LGCA) was designed to enhance the interaction of semantic information from LDM and global information from multi-head self-attention operations. Finally, a hierarchical metric was proposed to evaluate the similarity of image pairs in both the original pixel space and latent-feature space, enhancing topology preservation while improving registration accuracy. Extensive experiments on four public 2D cardiac image datasets, two 3D image datasets, show that the proposed LDM-Morph framework outperformed existing state-of-the-art CNNs- and Transformers-based registration methods regarding accuracy with comparable topology preservation and computational efficiency. Our code is publicly available at: https://github.com/wujiong-hub/LDM-Morph . Tinsu Pan, Kuang Gong |
Pattern Recognit. | 3 |
| 2026 | Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image AnalysisabstractVision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spatial features with clinical text. We present Med3DVLM, a 3D VLM designed to address these challenges through three key innovations: (1) DCFormer, an efficient encoder that uses decomposed 3D convolutions to capture fine-grained spatial features at scale; (2) SigLIP, a contrastive learning strategy with pairwise sigmoid loss that improves image-text alignment without relying on large negative batches; and (3) a dual-stream MLP-Mixer projector that fuses low- and high-level image features with text embeddings for richer multi-modal representations. We evaluated our model on the M3D dataset, which includes radiology reports and VQA data for 120,084 3D medical images. The results show that Med3DVLM achieves superior performance on multiple benchmarks. For image-text retrieval, it reaches 61.00% R@1 on 2,000 samples, significantly outperforming the current state-of-the-art M3D-LaMed model (19.10%). For report generation, it achieves a METEOR score of 36.42% (vs. 14.38%). In open-ended visual question answering (VQA), it scores 36.76% METEOR (vs. 33.58%), and in closed-ended VQA, it achieves 79.95% accuracy (vs. 75.78%). These results demonstrate Med3DVLM's ability to bridge the gap between 3D imaging and language, enabling scalable, multi-task reasoning across clinical applications. Gorkem Can Ates, Kuang Gong, Wei Shao 0008 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | PET Image Reconstruction Using Deep Diffusion Image PriorabstractDiffusion models have shown great promise in medical image denoising and reconstruction, but their application to Positron Emission Tomography (PET) imaging remains limited by tracer-specific contrast variability and high computational demands. In this work, we proposed an anatomical prior-guided PET image reconstruction method based on diffusion models, inspired by the deep diffusion image prior (DDIP) framework. The proposed method alternated between diffusion sampling and model fine-tuning guided by the PET sinogram, enabling the reconstruction of high-quality images from various PET tracers using a score function pretrained on a dataset of another tracer. To improve computational efficiency, the half-quadratic splitting (HQS) algorithm was adopted to decouple network optimization from iterative PET reconstruction. The proposed method was evaluated using one simulation and two clinical datasets. For the simulation study, a model pretrained on [ ${}^{{18}}\text {F}$ ]FDG data was tested on [ ${}^{{18}}\text {F}$ ]FDG data and amyloid-negative PET data to assess out-of-distribution (OOD) performance. For the clinical-data validation, ten low-dose [ ${}^{{18}}\text {F}$ ]FDG datasets and one [ ${}^{{18}}\text {F}$ ]Florbetapir dataset were tested on a model pretrained on data from another tracer. Experiment results show that the proposed PET reconstruction method can generalize robustly across tracer distributions and scanner types, providing an efficient and versatile reconstruction framework for low-dose PET imaging. Fumio Hashimoto, Kuang Gong |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Fast-DDPM: Fast Denoising Diffusion Probabilistic Models for Medical Image-to-Image GenerationabstractDenoising diffusion probabilistic models (DDPMs) have achieved unprecedented success in computer vision. However, they remain underutilized in medical imaging, a field crucial for disease diagnosis and treatment planning. This is primarily due to the high computational cost associated with the use of large number of time steps (e.g., 1,000) in diffusion processes. Training a diffusion model on medical images typically takes days to weeks, while sampling each image volume takes minutes to hours. To address this challenge, we introduce Fast-DDPM, a simple yet effective approach capable of simultaneously improving training speed, sampling speed, and generation quality. Unlike DDPM, which trains the image denoiser across 1,000 time steps, Fast-DDPM trains and samples using only 10 time steps. The key to our method lies in aligning the training and sampling procedures to optimize time-step utilization. Specifically, we introduced two efficient noise schedulers with 10 time steps: one with uniform time step sampling and another with non-uniform sampling. We evaluated Fast-DDPM across three medical image-to-image generation tasks: multi-image super-resolution, image denoising, and image-to-image translation. Fast-DDPM outperformed DDPM and current state-of-the-art methods based on convolutional networks and generative adversarial networks in all tasks. Additionally, Fast-DDPM reduced the training time to 0.2× and the sampling time to 0.01× compared to DDPM. Muhammad Imran 0013, Yuyin Zhou, Muxuan Liang, Kuang Gong, Wei Shao 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | PET Image Denoising Based on 3D Denoising Diffusion Probabilistic Model: Evaluations on Total-Body Datasets
Boxiao Yu, Savas Ozdemir, Yafei Dong, Wei Shao 0008, Kuangyu Shi, Kuang Gong |
MICCAI (7) | 6 |
| 2024 | Spach Transformer: Spatial and Channel-Wise Transformer Based on Local and Global Self-Attentions for PET Image DenoisingabstractPosition emission tomography (PET) is widely used in clinics and research due to its quantitative merits and high sensitivity, but suffers from low signal-to-noise ratio (SNR). Recently convolutional neural networks (CNNs) have been widely used to improve PET image quality. Though successful and efficient in local feature extraction, CNN cannot capture long-range dependencies well due to its limited receptive field. Global multi-head self-attention (MSA) is a popular approach to capture long-range information. However, the calculation of global MSA for 3D images has high computational costs. In this work, we proposed an efficient spatial and channel-wise encoder-decoder transformer, Spach Transformer, that can leverage spatial and channel information based on local and global MSAs. Experiments based on datasets of different PET tracers, i.e., 18F-FDG, 18F-ACBC, 18F-DCFPyL, and 68Ga-DOTATATE, were conducted to evaluate the proposed framework. Quantitative results show that the proposed Spach Transformer framework outperforms state-of-the-art deep learning architectures. Se-In Jang, Tinsu Pan, Pedram Heidari, Junyu Chen 0002, Quanzheng Li, Kuang Gong |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Anatomically Guided PET Image Reconstruction Using Conditional Weakly-Supervised Multi-Task Learning Integrating Self-AttentionabstractTo address the lack of high-quality training labels in positron emission tomography (PET) imaging, weakly-supervised reconstruction methods that generate network-based mappings between prior images and noisy targets have been developed. However, the learned model has an intrinsic variance proportional to the average variance of the target image. To suppress noise and improve the accuracy and generalizability of the learned model, we propose a conditional weakly-supervised multi-task learning (MTL) strategy, in which an auxiliary task is introduced serving as an anatomical regularizer for the PET reconstruction main task. In the proposed MTL approach, we devise a novel multi-channel self-attention (MCSA) module that helps learn an optimal combination of shared and task-specific features by capturing both local and global channel-spatial dependencies. The proposed reconstruction method was evaluated on NEMA phantom PET datasets acquired at different positions in a PET/CT scanner and 26 clinical whole-body PET datasets. The phantom results demonstrate that our method outperforms state-of-the-art learning-free and weakly-supervised approaches obtaining the best noise/contrast tradeoff with a significant noise reduction of approximately 50.0% relative to the maximum likelihood (ML) reconstruction. The patient study results demonstrate that our method achieves the largest noise reductions of 67.3% and 35.5% in the liver and lung, respectively, as well as consistently small biases in 8 tumors with various volumes and intensities. In addition, network visualization reveals that adding the auxiliary task introduces more anatomical information into PET reconstruction than adding only the anatomical loss, and the developed MCSA can abstract features and retain PET image details. Bao Yang, Kuang Gong, Huafeng Liu 0003, Quanzheng Li, Wentao Zhu 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Neural KEM: A Kernel Method With Deep Coefficient Prior for PET Image ReconstructionabstractImage reconstruction of low-count positron emission tomography (PET) data is challenging. Kernel methods address the challenge by incorporating image prior information in the forward model of iterative PET image reconstruction. The kernelized expectation-maximization (KEM) algorithm has been developed and demonstrated to be effective and easy to implement. A common approach for a further improvement of the kernel method would be adding an explicit regularization, which however leads to a complex optimization problem. In this paper, we propose an implicit regularization for the kernel method by using a deep coefficient prior, which represents the kernel coefficient image in the PET forward model using a convolutional neural-network. To solve the maximum-likelihood neural network-based reconstruction problem, we apply the principle of optimization transfer to derive a neural KEM algorithm. Each iteration of the algorithm consists of two separate steps: a KEM step for image update from the projection data and a deep-learning step in the image domain for updating the kernel coefficient image using the neural network. This optimization algorithm is guaranteed to monotonically increase the data likelihood. The results from computer simulations and real patient data have demonstrated that the neural KEM can outperform existing KEM and deep image prior methods. Kuang Gong, Ramsey Derek Badawi, Edward J. Kim 0002, Jinyi Qi, Guobao Wang |
IEEE Trans. Medical Imaging | 2 |
| 2022 | PET Denoising and Uncertainty Estimation Based on NVAE Model Using Quantile Regression Loss
Jianan Cui, Yutong Xie 0004, Anand A. Joshi, Kuang Gong, Kyung Sang Kim, Young-Don Son, Jong Hoon Kim, Richard M. Leahy, Huafeng Liu 0003, Quanzheng Li |
MICCAI (4) | 4 |
| 2022 | Unsupervised PET logan parametric image estimation using conditional deep image prior
Jianan Cui, Kuang Gong, Kyung Sang Kim, Huafeng Liu 0003, Quanzheng Li |
Medical Image Anal. | 2 |
| 2022 | Direct Reconstruction of Linear Parametric Images From Dynamic PET Using Nonlocal Deep Image PriorabstractDirect reconstruction methods have been developed to estimate parametric images directly from the measured PET sinograms by combining the PET imaging model and tracer kinetics in an integrated framework. Due to limited counts received, signal-to-noise-ratio (SNR) and resolution of parametric images produced by direct reconstruction frameworks are still limited. Recently supervised deep learning methods have been successfully applied to medical imaging denoising/reconstruction when large number of high-quality training labels are available. For static PET imaging, high-quality training labels can be acquired by extending the scanning time. However, this is not feasible for dynamic PET imaging, where the scanning time is already long enough. In this work, we proposed an unsupervised deep learning framework for direct parametric reconstruction from dynamic PET, which was tested on the Patlak model and the relative equilibrium Logan model. The training objective function was based on the PET statistical model. The patient’s anatomical prior image, which is readily available from PET/CT or PET/MR scans, was supplied as the network input to provide a manifold constraint, and also utilized to construct a kernel layer to perform non-local feature denoising. The linear kinetic model was embedded in the network structure as a${1} \times {1} \times {1}$convolution layer. Evaluations based on dynamic datasets of18F-FDG and11C-PiB tracers show that the proposed framework can outperform the traditional and the kernel method-based direct reconstruction methods. Kuang Gong, Ciprian Catana, Jinyi Qi, Quanzheng Li |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Clinically Translatable Direct Patlak Reconstruction from Dynamic PET with Motion Correction Using Convolutional Neural Network
Nuobei Xie, Kuang Gong, ZhiXing Qin, Jianan Cui, Zhifang Wu, Huafeng Liu 0003, Quanzheng Li |
MICCAI (7) | 2 |
| 2020 | Machine Learning in PET: From Photon Detection to Quantitative Image ReconstructionabstractMachine learning has found unique applications in nuclear medicine from photon detection to quantitative image reconstruction. While there have been impressive strides in detector development for time-of-flight positron emission tomography, most detectors still make use of simple signal processing methods to extract the time and position information from the detector signals. Now with the availability of fast waveform digitizers, machine learning techniques have been applied to estimate the position and arrival time of high-energy photons. In quantitative image reconstruction, machine learning has been used to estimate various corrections factors, including scattered events and attenuation images, as well as to reduce statistical noise in reconstructed images. Here machine learning either provides a faster alternative to an existing time-consuming computation, such as in the case of scatter estimation, or creates a data-driven approach to map an implicitly defined function, such as in the case of estimating the attenuation map for PET/MR scans. In this article, we will review the abovementioned applications of machine learning in nuclear medicine. Kuang Gong, Eric Berg, Simon R. Cherry, Jinyi Qi |
Proc. IEEE | 1 |
| 2020 | Severity and Consolidation Quantification of COVID-19 From CT Images Using Deep Learning Based on Hybrid Weak LabelsabstractEarly and accurate diagnosis of Coronavirus disease (COVID-19) is essential for patient isolation and contact tracing so that the spread of infection can be limited. Computed tomography (CT) can provide important information in COVID-19, especially for patients with moderate to severe disease as well as those with worsening cardiopulmonary status. As an automatic tool, deep learning methods can be utilized to perform semantic segmentation of affected lung regions, which is important to establish disease severity and prognosis prediction. Both the extent and type of pulmonary opacities help assess disease severity. However, manually pixel-level multi-class labelling is time-consuming, subjective, and non-quantitative. In this article, we proposed a hybrid weak label-based deep learning method that utilize both the manually annotated pulmonary opacities from COVID-19 pneumonia and the patient-level disease-type information available from the clinical report. A UNet was firstly trained with semantic labels to segment the total infected region. It was used to initialize another UNet, which was trained to segment the consolidations with patient-level information using the Expectation-Maximization (EM) algorithm. To demonstrate the performance of the proposed method, multi-institutional CT datasets from Iran, Italy, South Korea, and the United States were utilized. Results show that our proposed method can predict the infected regions as well as the consolidation regions with good correlation to human annotation. Dufan Wu, Kuang Gong, Chiara Daniela Arru, Fatemeh Homayounieh, Bernardo Bizzo, Varun Buch, Hui Ren 0001, Kyung Sang Kim, Nir Neumark, Nuobei Xie, Won Young Tak, Soo Young Park, Yu Rim Lee, Min Kyu Kang, Jung Gil Park, Alessandro Carriero, Luca Saba, Mahsa Masjedi, Hamidreza Talari, Rosa Babaei, Hadi Karimi Mobin, Shadi Ebrahimian, Ittai Dayan, Mannudeep K. Kalra, Quanzheng Li |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Consensus Neural Network for Medical Imaging Denoising with Only Noisy Training Samples
Dufan Wu, Kuang Gong, Kyung Sang Kim, Xiang Li 0001, Quanzheng Li |
MICCAI (4) | 2 |
| 2019 | PET Image Reconstruction Using Deep Image PriorabstractRecently, deep neural networks have been widely and successfully applied in computer vision tasks and have attracted growing interest in medical imaging. One barrier for the application of deep neural networks to medical imaging is the need for large amounts of prior training pairs, which is not always feasible in clinical practice. This is especially true for medical image reconstruction problems, where raw data are needed. Inspired by the deep image prior framework, in this paper, we proposed a personalized network training method where no prior training pairs are needed, but only the patient's own prior information. The network is updated during the iterative reconstruction process using the patient-specific prior information and measured data. We formulated the maximum-likelihood estimation as a constrained optimization problem and solved it using the alternating direction method of multipliers algorithm. Magnetic resonance imaging guided positron emission tomography reconstruction was employed as an example to demonstrate the effectiveness of the proposed framework. Quantification results based on simulation and real data show that the proposed reconstruction framework can outperform Gaussian post-smoothing and anatomically guided reconstructions using the kernel method or the neural-network penalty. Kuang Gong, Ciprian Catana, Jinyi Qi, Quanzheng Li |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Iterative PET Image Reconstruction Using Convolutional Neural Network RepresentationabstractPET image reconstruction is challenging due to the ill-poseness of the inverse problem and limited number of detected photons. Recently, the deep neural networks have been widely and successfully used in computer vision tasks and attracted growing interests in medical imaging. In this paper, we trained a deep residual convolutional neural network to improve PET image quality by using the existing inter-patient information. An innovative feature of the proposed method is that we embed the neural network in the iterative reconstruction framework for image representation, rather than using it as a post-processing tool. We formulate the objective function as a constrained optimization problem and solve it using the alternating direction method of multipliers algorithm. Both simulation data and hybrid real data are used to evaluate the proposed method. Quantification results show that our proposed iterative neural network method can outperform the neural network denoising and conventional penalized maximum likelihood methods. Kuang Gong, Jiahui Guan, Kyung Sang Kim, Xuezhu Zhang, Jaewon Yang, Youngho Seo, Georges El Fakhri, Jinyi Qi, Quanzheng Li |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Direct Patlak Reconstruction From Dynamic PET Data Using the Kernel Method With MRI Information Based on Structural SimilarityabstractPositron emission tomography (PET) is a functional imaging modality widely used in oncology, cardiology, and neuroscience. It is highly sensitive, but suffers from relatively poor spatial resolution, as compared with anatomical imaging modalities, such as magnetic resonance imaging (MRI). With the recent development of combined PET/MR systems, we can improve the PET image quality by incorporating MR information into image reconstruction. Previously, kernel learning has been successfully embedded into static and dynamic PET image reconstruction using either PET temporal or MRI information. Here, we combine both PET temporal and MRI information adaptively to improve the quality of direct Patlak reconstruction. We examined different approaches to combine the PET and MRI information in kernel learning to address the issue of potential mismatches between MRI and PET signals. Computer simulations and hybrid real-patient data acquired on a simultaneous PET/MR scanner were used to evaluate the proposed methods. Results show that the method that combines PET temporal information and MRI spatial information adaptively based on the structure similarity index has the best performance in terms of noise reduction and resolution improvement. Kuang Gong, Jinxiu Cheng-Liao, Guobao Wang, Kevin T. Chen, Ciprian Catana, Jinyi Qi |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Corrections to "Direct Patlak Reconstruction From Dynamic PET Data Using the Kernel Method With MRI Information Based on Structural Similarity"abstractIn the above paper[1], there are typos inAlgorithm 1table. The correct version ofAlgorithm 1is given below. Kuang Gong, Jinxiu Cheng-Liao, Guobao Wang, Kevin T. Chen, Ciprian Catana, Jinyi Qi |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Penalized PET Reconstruction Using Deep Learning Prior and Local Linear FittingabstractMotivated by the great potential of deep learning in medical imaging, we propose an iterative positron emission tomography reconstruction framework using a deep learning-based prior. We utilized the denoising convolutional neural network (DnCNN) method and trained the network using full-dose images as the ground truth and low dose images reconstructed from downsampled data by Poisson thinning as input. Since most published deep networks are trained at a predetermined noise level, the noise level disparity of training and testing data is a major problem for their applicability as a generalized prior. In particular, the noise level significantly changes in each iteration, which can potentially degrade the overall performance of iterative reconstruction. Due to insufficient existing studies, we conducted simulations and evaluated the degradation of performance at various noise conditions. Our findings indicated that DnCNN produces additional bias induced by the disparity of noise levels. To address this issue, we propose a local linear fitting function incorporated with the DnCNN prior to improve the image quality by preventing unwanted bias. We demonstrate that the resultant method is robust against noise level disparities despite the network being trained at a predetermined noise level. By means of bias and standard deviation studies via both simulations and clinical experiments, we show that the proposed method outperforms conventional methods based on total variation and non-local means penalties. We thereby confirm that the proposed method improves the reconstruction result both quantitatively and qualitatively. Kyung Sang Kim, Dufan Wu, Kuang Gong, Joyita Dutta, Jong Hoon Kim, Young-Don Son, Hang-Keun Kim, Georges El Fakhri, Quanzheng Li |
IEEE Trans. Medical Imaging | 3 |
| 2017 | Sinogram Blurring Matrix Estimation From Point Sources Measurements With Rank-One Approximation for Fully 3-D PETabstractAn accurate system matrix is essential in positron emission tomography (PET) for reconstructing high quality images. To reduce storage size and image reconstruction time, we factor the system matrix into a product of a geometry projection matrix and a sinogram blurring matrix. The geometric projection matrix is computed analytically and the sinogram blurring matrix is estimated from point source measurements. Previously, we have estimated a 2-D blurring matrix for a preclinical PET scanner. The 2-D blurring matrix only considers blurring effects within a transaxial sinogram and does not compensate for inter-sinogram blurring effects. For PET scanners with a long axial field of view, inter-sinogram blurring can be a major problem influencing the image quality in the axial direction. Hence, the estimation of a 4-D blurring matrix is desirable to further improve the image quality. The 4-D blurring matrix estimation is an ill-conditioned problem due to the large number of unknowns. Here, we propose a rank-one approximation for each blurring kernel image formed by a row vector of the sinogram blurring matrix to improve the stability of the 4-D blurring matrix estimation. The proposed method is applied to the simulated data as well as the real data obtained from an Inveon microPET scanner. The results show that the newly estimated 4-D blurring matrix can improve the image quality over those obtained with a 2-D blurring matrix and requires less point source scans to achieve similar image quality compared with an unconstrained 4-D blurring matrix estimation. Kuang Gong, Michel Tohme, Martin Judenhofer, Yongfeng Yang, Jinyi Qi |
IEEE Trans. Medical Imaging | 1 |