EDBT 2026 Demo / reviewers in the wild / expert
Pujin Cheng
dblp:250/6030
· DBLP profile ↗
22ranked-venue papers
3as first author
22since 2021 · last 2026
0009-0000-4844-5146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hessian-Aware Zeroth-Order Optimization for Quantized Large Language Models
Zhangchi Wang, Yidu Wu, Yijin Huang, Pujin Cheng, Qinghai Guo, Xiaoying Tang 0001 |
ICIC (5) | 4 |
| 2026 | MIRAGE: Medical image-text pre-training for robustness against noisy environments
Pujin Cheng, Yijin Huang, Li Lin 0006, Junyan Lyu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
Medical Image Anal. | 1 |
| 2026 | AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware TransformerabstractIn the era of foundation models, achieving a unified understanding of different dynamic objects through a single network has the potential to empower stronger spatial intelligence. Moreover, accurate estimation of animal pose and shape across diverse species is essential for quantitative analysis in biological research. However, this topic remains underexplored due to the limited network capacity of previous methods and the scarcity of comprehensive multi-species datasets. To address these limitations, we introduce AniMer+, an extended version of our scalable AniMer framework. In this paper, we focus on a unified approach for reconstructing mammals (mammalia) and birds (aves). A key innovation of AniMer+ is its high-capacity, family-aware Vision Transformer (ViT) incorporating a Mixture-of-Experts (MoE) design. Its architecture partitions network layers into taxa-specific components (for mammalia and aves) and taxa-shared components, enabling efficient learning of both distinct and common anatomical features within a single model. To overcome the critical shortage of 3D training data, especially for birds, we introduce a diffusion-based conditional image generation pipeline. This pipeline produces two large-scale synthetic datasets: CtrlAni3D for quadrupeds (about 10 k images with pixel-aligned SMAL labels) and CtrlAVES3D (about 7 k images with pixel-aligned AVES labels). To note, CtrlAVES3D is the first large-scale, 3D-annotated dataset for birds, which is crucial for resolving single-view depth ambiguities. Trained on an aggregated collection of 41.3 k mammalian and 12.4 k avian images (combining real and synthetic data), our method demonstrates superior performance over existing approaches across a wide range of benchmarks, including the challenging out-of-domain Animal Kingdom dataset. Ablation studies confirm the effectiveness of both our novel network architecture and the generated synthetic datasets in enhancing real-world application performance. Liang An 0001, Jin Lyu, Li Lin 0006, Pujin Cheng, Yebin Liu, Xiaoying Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | SynLiTS: Phase Prompting-Driven Diffusion Synthesis and Context-Aware Fusion for Unaligned Liver Tumor SegmentationabstractGiven the rich and complementary information contained in multi-phase CT images (CTs), they play an indispensable role in liver cancer diagnosis and prognosis, wherein an important prerequisite is liver tumor segmentation. However, spatial misalignments across phases and the limited availability of high-quality multi-phase CT datasets significantly hinder the performance of liver tumor segmentation. To tackle these challenges, we here propose SynLiTS, a novel multi-phase liver tumor segmentation framework. Its core idea is to synthesize multi-phase CTs with strictly-aligned liver tumors based on pseudo-normal multi-phase CTs. Specifically, an FFC-based Inpainter is first designed to generate pseudo-normal CTs by reconstructing dilated liver tumors. The pseudo-normal CTs and randomly generated tumor masks are then combined via a phase prompting-driven diffusion model to synthesize multi-phase liver tumor CTs with diverse tumor characteristics. In this way, multi-phase CTs with perfectly-aligned liver tumor labels are obtained. We also construct a real multi-phase liver tumor dataset, named MPLiTS. Finally, the synthesized and real multi-phase CTs are used to train a liver tumor segmentation model, which incorporates a context-aware fusion module to effectively learn and integrate multi-phase information. SynLiTS is evaluated on both internal and external datasets, and the results show that it outperforms state-of-the-art methods by large margins. Code will be released at https://github.com/Chyiun/SynLiTS. Li Lin 0006, ZhiCheng Jin, Pujin Cheng, JianJian Chen, HaiDong Zhu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | UNiFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial FusionabstractRecently, Multi-Contrast MR Reconstruction (MCMR) has emerged as a hot research topic that leverages high-quality auxiliary modalities to reconstruct undersampled target modalities of interest. However, existing methods often struggle to generalize across different k-space undersampling patterns, requiring the training of a separate model for each specific pattern, which limits their practical applicability. To address this challenge, we propose UniFS, a Unified Frequency-Spatial Fusion model designed to handle multiple k-space undersampling patterns for MCMR tasks without any need for retraining. UniFS integrates three key modules: a Cross-Modal Frequency Fusion module, an Adaptive Mask-Based Prompt Learning module, and a Dual-Branch Complementary Refinement module. These modules work together to extract domain-invariant features from diverse k-space undersampling patterns while dynamically adapt to their own variations. Another limitation of existing MCMR methods is their tendency to focus solely on spatial information while neglect frequency characteristics, or extract only shallow frequency features, thus failing to fully leverage complementary cross-modal frequency information. To relieve this issue, UniFS introduces an adaptive prompt-guided frequency fusion module for k-space learning, significantly enhancing the model's generalization performance. We evaluate our model on the BraTS and HCP datasets with various k-space undersampling patterns and acceleration factors, including previously unseen patterns, to comprehensively assess UniFS's generalizability. Experimental results across multiple scenarios demonstrate that UniFS achieves state-of-the-art performance. Our code is available at https://github.com/LIKP0/UniFS. Yiwei Ren, Kai Pan, Dong Wei 0004, Pujin Cheng, Xian Wu 0001, Xiaoying Tang 0001 |
BIBM | 5 |
| 2025 | AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerabstractQuantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-ofdistribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. Project page: https://luoxue-star.github.io/AniMer_project_page/. Jin Lyu, Yi Gu 0005, Li Lin 0006, Pujin Cheng, Yebin Liu, Xiaoying Tang 0001, Liang An 0001 |
CVPR | 5 |
| 2025 | LF-SynthSeg: Label-Free Brain Tissue-Assisted Tumor Synthesis and SegmentationabstractUnsupervised brain tumor segmentation is pivotal in realms of disease diagnosis, surgical planning, and treatment response monitoring, with the distinct advantage of obviating the need for labeled data. Traditional methodologies in this domain, however, often fall short in fully capitalizing on the extensive prior knowledge of brain tissue, typically approaching the task merely as an anomaly detection challenge. In our research, we present an innovative strategy that effectively integrates brain tissues' prior knowledge into both the synthesis and segmentation of brain tumor from T2-weighted Magnetic Resonance Imaging scans. Central to our method is the tumor synthesis mechanism, employing randomly generated ellipsoids in conjunction with the intensity profiles of brain tissues. This methodology not only fosters a significant degree of variation in the tumor presentations within the synthesized images but also facilitates the creation of an essentially unlimited pool of abnormal T2-weighted images. These synthetic images closely replicate the characteristics of real tumor-bearing scans. Our training protocol extends beyond mere tumor segmentation; it also encompasses the segmentation of brain tissues, thereby directing the network's attention to the boundary relationship between brain tumor and brain tissue, thus improving the robustness of our method. We evaluate our approach across five widely recognized public datasets (BRATS 2019, BRATS 2020, BRATS 2021, PED and SSA), and the results show that our method outperforms state-of-the-art unsupervised tumor segmentation methods by large margins. Moreover, the proposed method achieves more than 92 of the fully supervised performance on the same testing datasets. Pengxiao Xu, Junyan Lyu, Li Lin 0006, Pujin Cheng, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | FedLPPA: Learning Personalized Prompt and Aggregation for Federated Weakly-Supervised Medical Image SegmentationabstractFederated learning (FL) effectively mitigates the data silo challenge brought about by policies and privacy concerns, implicitly harnessing more data for deep model training. However, traditional centralized FL models grapple with diverse multi-center data, especially in the face of significant data heterogeneity, notably in medical contexts. In the realm of medical image segmentation, the growing imperative to curtail annotation costs has amplified the importance of weakly-supervised techniques which utilize sparse annotations such as points, scribbles, etc. A pragmatic FL paradigm shall accommodate diverse annotation formats across different sites, which research topic remains under-investigated. In such context, we propose a novel personalized FL framework with learnable prompt and aggregation (FedLPPA) to uniformly leverage heterogeneous weak supervision for medical image segmentation. In FedLPPA, a learnable universal knowledge prompt is maintained, complemented by multiple learnable personalized data distribution prompts and prompts representing the supervision sparsity. Integrated with sample features through a dual-attention mechanism, those prompts empower each local task decoder to adeptly adjust to both the local distribution and the supervision form. Concurrently, a dual-decoder strategy, predicated on prompt similarity, is introduced for enhancing the generation of pseudo-labels in weakly-supervised learning, alleviating overfitting and noise accumulation inherent to local data, while an adaptable aggregation method is employed to customize the task decoder on a parameter-wise basis. Extensive experiments on four distinct medical image segmentation tasks involving different modalities underscore the superiority of FedLPPA, with its efficacy closely parallels that of fully supervised centralized training. Our code and data will be available at https://github.com/llmir/FedLPPA. Li Lin 0006, Yixiang Liu, Jiewei Wu, Pujin Cheng, Zhiyuan Cai, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Boosting Memory Efficiency in Transfer Learning for High-Resolution Medical Image ClassificationabstractThe success of large-scale pretrained models has established fine-tuning as a standard method for achieving significant improvements in downstream tasks. However, fine-tuning the entire parameter set of a pretrained model is costly. Parameter-efficient transfer learning (PETL) has recently emerged as a cost-effective alternative for adapting pretrained models to downstream tasks. Despite its advantages, the increasing model size and input resolution present challenges for PETL, as the training memory consumption is not reduced as effectively as the parameter usage. In this article, we introduce fine-grained prompt tuning plus (FPT+), a PETL method designed for high-resolution medical image classification, which significantly reduces the training memory consumption compared to other PETL methods. FPT+ performs transfer learning by training a lightweight side network and accessing pretrained knowledge from a large pretrained model (LPM) through fine-grained prompts and fusion modules. Specifically, we freeze the LPM of interest and construct a learnable lightweight side network. The frozen LPM processes high-resolution images to extract fine-grained features, while the side network employs corresponding downsampled low-resolution images to minimize memory usage. To enable the side network to leverage pretrained knowledge, we propose fine-grained prompts and fusion modules, which collaborate to summarize information through the LPM's intermediate activations. We evaluate FPT+ on eight medical image datasets of varying sizes, modalities, and complexities. Experimental results demonstrate that FPT+ outperforms other PETL methods, using only 1.03% of the learnable parameters and 3.18% of the memory required for fine-tuning an entire ViT-B model. Our code is available https://github.com/YijinHuang/FPT. Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Joint Super-Resolution and Modality Translation Network for Multi-Contrast Arbitrary-Scale Isotropic MRI ReconstructionabstractDue to time and cost limitations, Magnetic Resonance (MR) imaging often employs anisotropic scanning with large slice spacing and thickness. This causes blurring in views perpendicular to the slices, which adversely affects clinical diagnosis and research. Taking into account the complementary information from the reference modality, deep learning (DL) based multi-contrast methods have become a focal point of research. These methods aim to reconstruct the isotropic target MR image with the auxiliary high-resolution (HR) reference modality. However, most of the methods primarily concentrate on the structural restoration of the target low-resolution (LR) image, neglecting the crucial aspect that the coexisting structural and modality differences between target and reference modalities can impede effective restoration. Additionally, these methods are designed for a fixed upsampling scale, not accounting for the practical scenario of varying slice thickness. In this work, we propose a joint Super-resolution and Modality translation network (SMNet) for multi-contrast arbitrary-scale isotropic MRI reconstruction. The modality translation branch includes the Modality-Specific-Augmented Alignment (MSAA) block, which eliminates modality distribution disparities and enhances modality-specific regions on the reference feature before fusion. And the super-resolution branch employs the Reliability-based Spatial Fusion (RSF) block for the structural restoration of the target LR feature using a reliability prior. The outputs from these two branches are then ensembled to obtain the final reconstructed result. Extensive experiments on both a private dataset and the Brasts2021 dataset demonstrate the effectiveness and generalizability of the proposed method. Our code is available at https://github.com/11710615/smnet. Kai Pan, Li Lin 0006, Pujin Cheng, Junyan Lyu, Xiaoying Tang 0001 |
BIBM | 3 |
| 2024 | Fusion Side Tuning: A Parameter and Memory Efficient Fine-tuning Method for High-resolution Medical Image ClassificationabstractParameter-efficient fine-tuning (PEFT) has been proposed as a cost-effective approach for transferring large-scale pre-trained models (LPMs) to downstream tasks, mitigating the high costs associated with updating all parameters of LPMs. However, current PEFT methods encounter the challenge that GPU memory usage during training is not reduced as effectively as parameter usage. In this paper, we propose Fusion Side Tuning (FST), a novel memory-efficient and parameter-efficient fine-tuning method. FST significantly reduces GPU memory consumption during training, particularly at high resolutions. To achieve this, we freeze the backbone LPM and construct a learnable side fusion network that takes intermediate features from the backbone as input. The side fusion network consists of a sequence of fusion modules, which enable it to leverage the knowledge embedded in the intermediate features. Additionally, we employ an important token selection mechanism to further reduce training costs and memory requirements. We evaluate FST on eight medical image datasets of varying modalities and sizes. Experimental results demonstrate that FST outperforms existing PEFT methods, utilizing only 2% of the learnable parameters and 20% of the GPU memory required for full fine-tuning of a ViT-B encoder with an input resolution of 512 × 512. Zhangchi Wang, Yijin Huang, Yidu Wu, Pujin Cheng, Li Lin 0006, Qinghai Guo, Xiaoying Tang 0001 |
BIBM | 4 |
| 2024 | Fine-Grained Prompt Tuning: A Parameter and Memory Efficient Transfer Learning Method for High-Resolution Medical Image Classification
Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
MICCAI (12) | 2 |
| 2024 | LDDMM-Face: Large deformation diffeomorphic metric learning for cross-annotation face alignment
Junyan Lyu, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
Pattern Recognit. | 3 |
| 2024 | SSiT: Saliency-Guided Self-Supervised Image Transformer for Diabetic Retinopathy GradingabstractSelf-supervised Learning (SSL) has been widely applied to learn image representations through exploiting unlabeled images. However, it has not been fully explored in the medical image analysis field. In this work, Saliency-guided Self-Supervised image Transformer (SSiT) is proposed for Diabetic Retinopathy (DR) grading from fundus images. We novelly introduce saliency maps into SSL, with a goal of guiding self-supervised pre-training with domain-specific prior knowledge. Specifically, two saliency-guided learning tasks are employed in SSiT: 1) Saliency-guided contrastive learning is conducted based on the momentum contrast, wherein fundus images' saliency maps are utilized to remove trivial patches from the input sequences of the momentum-updated key encoder. Thus, the key encoder is constrained to provide target representations focusing on salient regions, guiding the query encoder to capture salient features. 2) The query encoder is trained to predict the saliency segmentation, encouraging the preservation of fine-grained information in the learned representations. To assess our proposed method, four publicly-accessible fundus image datasets are adopted. One dataset is employed for pre-training, while the three others are used to evaluate the pre-trained models' performance on downstream DR grading. The proposed SSiT significantly outperforms other representative state-of-the-art SSL methods on all downstream datasets and under various evaluation settings. For example, SSiT achieves a Kappa score of 81.88% on the DDR dataset under fine-tuning evaluation, outperforming all other ViT-based SSL methods by at least 9.48%. Yijin Huang, Junyan Lyu, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Uni4Eye++: A General Masked Image Modeling Multi-Modal Pre-Training Framework for Ophthalmic Image Classification and SegmentationabstractA large-scale labeled dataset is a key factor for the success of supervised deep learning in most ophthalmic image analysis scenarios. However, limited annotated data is very common in ophthalmic image analysis, since manual annotation is time-consuming and labor-intensive. Self-supervised learning (SSL) methods bring huge opportunities for better utilizing unlabeled data, as they do not require massive annotations. To utilize as many unlabeled ophthalmic images as possible, it is necessary to break the dimension barrier, simultaneously making use of both 2D and 3D images as well as alleviating the issue of catastrophic forgetting. In this paper, we propose a universal self-supervised Transformer framework named Uni4Eye++ to discover the intrinsic image characteristic and capture domain-specific feature embedding in ophthalmic images. Uni4Eye++ can serve as a global feature extractor, which builds its basis on a Masked Image Modeling task with a Vision Transformer architecture. On the basis of our previous work Uni4Eye, we further employ an image entropy guided masking strategy to reconstruct more-informative patches and a dynamic head generator module to alleviate modality confusion. We evaluate the performance of our pre-trained Uni4Eye++ encoder by fine-tuning it on multiple downstream ophthalmic image classification and segmentation tasks. The superiority of Uni4Eye++ is successfully established through comparisons to other state-of-the-art SSL pre-training methods. Our code is available at https://github.com/Davidczy/Uni4Eye++. Zhiyuan Cai, Li Lin 0006, Huaqing He, Pujin Cheng, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsabstractContrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports. In contrast to standard global multi-modality alignment methods, we employ a local alignment module for fine-grained representation. Furthermore, a cross-modality conditional reconstruction module is designed to interchange information across modalities in the training phase by reconstructing masked images and reports. For reconstructing long reports, a sentence-wise prototype memory bank is constructed, enabling the network to focus on low-level localized visual and high-level clinical linguistic features. Additionally, a non-auto-regressive generation paradigm is proposed for reconstructing non-sequential reports. Experimental results on five downstream tasks, including supervised classification, zero-shot classification, image-to-text retrieval, semantic segmentation, and object detection, show the proposed method outperforms other state-of-the-art methods across multiple datasets and under different dataset size settings. The code is available at https://github.com/QtacierP/PRIOR. Pujin Cheng, Li Lin 0006, Junyan Lyu, Yijin Huang, Wenhan Luo, Xiaoying Tang 0001 |
ICCV | 1 |
| 2023 | YoloCurvSeg: You only label one noisy skeleton for vessel-style curvilinear structure segmentation
Li Lin 0006, Linkai Peng, Huaqing He, Pujin Cheng, Jiewei Wu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
Medical Image Anal. | 4 |
| 2022 | DS3-Net: Difficulty-Perceived Common-to-T1ce Semi-supervised Multimodal MRI Synthesis Network
Li Lin 0006, Pujin Cheng, Kai Pan, Xiaoying Tang 0001 |
MICCAI (6) | 3 |
| 2022 | AADG: Automatic Augmentation for Domain Generalization on Retinal Image SegmentationabstractConvolutional neural networks have been widely applied to medical image segmentation and have achieved considerable performance. However, the performance may be significantly affected by the domain gap between training data (source domain) and testing data (target domain). To address this issue, we propose a data manipulation based domain generalization method, called Automated Augmentation for Domain Generalization (AADG). Our AADG framework can effectively sample data augmentation policies that generate novel domains and diversify the training set from an appropriate search space. Specifically, we introduce a novel proxy task maximizing the diversity among multiple augmented novel domains as measured by the Sinkhorn distance in a unit sphere space, making automated augmentation tractable. Adversarial training and deep reinforcement learning are employed to efficiently search the objectives. Quantitative and qualitative experiments on 11 publicly-accessible fundus image datasets (four for retinal vessel segmentation, four for optic disc and cup (OD/OC) segmentation and three for retinal lesion segmentation) are comprehensively performed. Two OCTA datasets for retinal vasculature segmentation are further involved to validate cross-modality generalization. Our proposed AADG exhibits state-of-the-art generalization performance and outperforms existing approaches by considerable margins on retinal vessel, OD/OC and lesion segmentation tasks. The learned policies are empirically validated to be model-agnostic and can transfer well to other models. The source code is available at https://github.com/CRazorback/AADG. Junyan Lyu, Yijin Huang, Li Lin 0006, Pujin Cheng, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | I-SECRET: Importance-Guided Fundus Image Enhancement via Semi-supervised Contrastive Constraining
Pujin Cheng, Li Lin 0006, Yijin Huang, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (8) | 1 |
| 2021 | Lesion-Based Contrastive Learning for Diabetic Retinopathy Grading from Fundus Images
Yijin Huang, Li Lin 0006, Pujin Cheng, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (2) | 3 |
| 2021 | BSDA-Net: A Boundary Shape and Distance Aware Joint Learning Framework for Segmenting and Classifying OCTA Images
Li Lin 0006, Jiewei Wu, Yijin Huang, Junyan Lyu, Pujin Cheng, Xiaoying Tang 0001 |
MICCAI (8) | 6 |