Yingtai Li

dblp:195/9285 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0002-5145-2245ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2025 AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIP
abstract
Anomaly detection (AD) identifies outliers for applications like defect and lesion detection. While CLIP shows promise for zero-shot AD tasks due to its strong generalization capabilities, its inherent Anomaly-Unawareness leads to limited discrimination between normal and abnormal features. To address this problem, we propose Anomaly-Aware CLIP (AA-CLIP), which enhances CLIP's anomaly discrimination ability in both text and visual spaces while preserving its generalization capability. AA-CLIP is achieved through a straightforward yet effective two-stage approach: it first creates anomaly-aware text anchors to differentiate normal and abnormal semantics clearly, then aligns patch-level visual features with these anchors for precise anomaly localization. This two-stage strategy, with the help of residual adapters, gradually adapts CLIP in a controlled manner, achieving effective AD while maintaining CLIP's class knowledge. Extensive experiments validate AA-CLIP as a resource-efficient solution for zero-shot AD tasks, achieving state-of-the-art results in industrial and medical applications. The code is available at https://github.com/Mwxinnn/AA-CLIP.
Qingsong Yao, Fenghe Tang, Chenxu Wu, Yingtai Li, Rui Yan 0009, Zihang Jiang, Shaohua Kevin Zhou
CVPR6
2025 Dyna3DGR: 4D Cardiac Motion Tracking with Dynamic 3D Gaussian Representation
Xueming Fu, Yingtai Li, Zihang Jiang, Junhao Mei, Gaojun Teng, Shaohua Kevin Zhou
MICCAI (2)3
2025 More Performant and Scalable: Rethinking Contrastive Vision-Language Pre-training of Radiology in the LLM Era
Yingtai Li, Haoran Lai, Xiaoqian Zhou, Shuai Ming, Wei Wei 0006, Shaohua Kevin Zhou
MICCAI (7)1
2025 3DGR-CT: Sparse-view CT reconstruction with a 3D Gaussian representation
abstract
Sparse-view computed tomography (CT) reduces radiation exposure by acquiring fewer projections, making it a valuable tool in clinical scenarios where low-dose radiation is essential. However, this often results in increased noise and artifacts due to limited data. In this paper we propose a novel 3D Gaussian representation (3DGR) based method for sparse-view CT reconstruction. Inspired by recent success in novel view synthesis driven by 3D Gaussian splatting, we leverage the efficiency and expressiveness of 3D Gaussian representation as an alternative to implicit neural representation. To unleash the potential of 3DGR for CT imaging scenario, we propose two key innovations: (i) FBP-image-guided Guassian initialization and (ii) efficient integration with a differentiable CT projector. Extensive experiments and ablations on diverse datasets demonstrate the proposed 3DGR-CT consistently outperforms state-of-the-art counterpart methods, achieving higher reconstruction accuracy with faster convergence. Furthermore, we showcase the potential of 3DGR-CT for real-time physical simulation, which holds important clinical applications while challenging for implicit neural representations. Code available at: https://github.com/SigmaLDC/3DGR-CT.
Yingtai Li, Xueming Fu, Shang Zhao 0004, Ruiyang Jin, Shaohua Kevin Zhou
Medical Image Anal.1
2025 MambaMIM: Pre-training Mamba with state space token interpolation and its application to medical image segmentation
Fenghe Tang, Bingkun Nian, Yingtai Li, Zihang Jiang, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou
Medical Image Anal.3
2024 Taming Stable Diffusion for MRI Cross-Modality Translation
abstract
In this study, we explore using Stable Diffusion (SD) for unsupervised medical image-to-image translation. SD has shown remarkable performances in generating high-quality images and can be easily applied to generate custom contents by injecting standard plug-ins like LoRA, offering a promising solution to tackle the complexity caused by variations in imaging modalities, acquisition parameters, and body parts in medical imaging. However, We empirically find that existing pipelines designed for natural images fail to translate directly to medical images due to weak structural control and inappropriate color preservation. To address these issues, we propose a novel two-branch image translation pipeline. This pipeline decouples the generation of target image along the time axis and employs ControlNet to ensure precise structural preservation. Additionally, we customize SD to generate images of extreme brightness, a common feature in medical imaging. Our results on the BraTS dataset demonstrate that SD with task-specific plug-ins can generate high-quality medical images comparable to those generated by task-specific models. Since the development of these standard plug-ins can be easily done by clinicians without much knowledge of the underlying algorithm, such a mode holds the potential to significantly extend the use of medical image computing algorithms in the clinical environment.
Yingtai Li, Shaohua Kevin Zhou
BIBM1
2024 3DGR-CAR: Coronary Artery Reconstruction from Ultra-sparse 2D X-Ray Views with a 3D Gaussians Representation
Xueming Fu, Yingtai Li, Fenghe Tang, Jun Li 0103, Mingyue Zhao, Gaojun Teng, Shaohua Kevin Zhou
MICCAI (7)2
2020 AprilE: Attention with Pseudo Residual Connection for Knowledge Graph Embedding
abstract
Knowledge graph embedding maps entities and relations into low-dimensional vector space.However, it is still challenging for many existing methods to model diverse relational patterns, especially symmetric and antisymmetric relations.To address this issue, we propose a novel model, AprilE, which employs triple-level self-attention and pseudo residual connection to model relational patterns.The triple-level self-attention treats head entity, relation, and tail entity as a sequence and captures the dependency within a triple.At the same time the pseudo residual connection retains primitive semantic features.Furthermore, to deal with symmetric and antisymmetric relations, two schemas of score function are designed via a position-adaptive mechanism.Experimental results on public datasets demonstrate that our model can produce expressive knowledge embedding and significantly outperforms most of the state-of-the-art works.
Yuzhang Liu, Peng Wang 0004, Yingtai Li, Yizhan Shao, Zhongkai Xu
COLING3