Tan Pan

dblp:231/1885 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Tracing the Heart's Pathways: ECG Representation Learning from a Cardiac Conduction Perspective
abstract
The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key limitation: they focus on consistent patterns across leads and beats, overlooking the inherent differences in heartbeats rooted in cardiac conduction processes, while subtle but significant variations carry unique physiological signatures. Moreover, representation learning for ECG analysis should align with ECG diagnostic guidelines, which progress from individual heartbeats to single leads and ultimately to lead combinations. This sequential logic, however, is often neglected when applying pre-trained models to downstream tasks. To address these gaps, we propose CLEAR-HUG, a two-stage framework designed to capture subtle variations in cardiac conduction across leads while adhering to ECG diagnostic guidelines. In the first stage, we introduce an eSSL model termed Conduction-LEAd Reconstructor (CLEAR), which captures both specific variations and general commonalities across heartbeats. Treating each heartbeat as a distinct entity, CLEAR employs a simple yet effective sparse attention mechanism to reconstruct signals without interference from other heartbeats. In the second stage, we implement a Hierarchical lead-Unified Group head (HUG) for disease diagnosis, mirroring clinical workflow. Experimental results across six tasks show a 6.84% improvement, validating the effectiveness of CLEAR-HUG. This highlights its ability to enhance representations of cardiac conduction and align patterns with expert diagnostic guidelines.
Tan Pan, Yixuan Sun, Chen Jiang 0006, Qiong Gao, Xingmeng Zhang, Zhenqi Yang, Limei Han, Yixiu Liang, Kaiyu Guo
AAAI1
2025 Structure-Aware Semantic Discrepancy and Consistency for 3D Medical Image Self-Supervised Learning
Tan Pan, Zhaorui Tan, Kaiyu Guo, Dongli Xu, Weidi Xu, Chen Jiang 0006, Xin Guo 0010, Yuan Qi 0001
ICCV1
2025 Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang
ICCV3
2025 Improving Out-of-Distribution Detection via Dynamic Covariance Calibration
abstract
Out-of-Distribution (OOD) detection is essential for the trustworthiness of AI systems. Methods using prior information (i.e., subspace-based methods) have shown effective performance by extracting information geometry to detect OOD data with a more appropriate distance metric. However, these methods fail to address the geometry distorted by ill-distributed samples, due to the limitation of statically extracting information geometry from the training distribution. In this paper, we argue that the influence of ill-distributed samples can be corrected by dynamically adjusting the prior geometry in response to new data. Based on this insight, we propose a novel approach that dynamically updates the prior covariance matrix using real-time input features, refining its information. Specifically, we reduce the covariance along the direction of real-time input features and constrain adjustments to the residual space, thus preserving essential data characteristics and avoiding effects on unintended directions in the principal space. We evaluate our method on two pre-trained models for the CIFAR dataset and five pre-trained models for ImageNet-1k, including the self-supervised DINO model. Extensive experiments demonstrate that our approach significantly enhances OOD detection across various models. The code is released at https://github.com/workerbcd/ooddcc.
Kaiyu Guo, Zijian Wang 0009, Tan Pan, Brian C. Lovell, Mahsa Baktash
ICML3
2025 Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
abstract
The generalization ability of deep learning has been extensively studied in supervised settings, yet it remains less explored in unsupervised scenarios. Recently, the Unsupervised Domain Generalization (UDG) task has been proposed to enhance the generalization of models trained with prevalent unsupervised learning techniques, such as Self-Supervised Learning (SSL). UDG confronts the challenge of distinguishing semantics from variations without category labels. Although some recent methods have employed domain labels to tackle this issue, such domain labels are often unavailable in real-world contexts. In this paper, we address these limitations by formalizing UDG as the task of learning a Minimal Sufficient Semantic Representation: a representation that (i) preserves all semantic information shared across augmented views (sufficiency), and (ii) maximally removes information irrelevant to semantics (minimality). We theoretically ground these objectives from the perspective of information theory, demonstrating that optimizing representations to achieve sufficiency and minimality directly reduces out-of-distribution risk. Practically, we implement this optimization through Minimal-Sufficient UDG (MS-UDG), a learnable model by integrating (a) an InfoNCE-based objective to achieve sufficiency; (b) two complementary components to promote minimality: a novel semantic-variation disentanglement loss and a reconstruction-based mechanism for capturing adequate variation. Empirically, MS-UDG sets a new state-of-the-art on popular unsupervised domain-generalization benchmarks, consistently outperforming existing SSL and UDG methods, without category or domain labels during representation learning.
Tan Pan, Kaiyu Guo, Dongli Xu, Zhaorui Tan, Chen Jiang 0006, Deshu Chen, Xin Guo 0010, Brian C. Lovell, Limei Han, Mahsa Baktash
NeurIPS1
2025 EdgeSegDiff: Edge-Conditional Diffusion Model for Skin Lesion Segmentation
Tan Pan, Nan Mu
PRCV (14)1
2023 Boundary-aware Backward-Compatible Representation via Adversarial Learning in Image Retrieval
abstract
Image retrieval plays an important role in the Internet world. Usually, the core parts of mainstream visual retrieval systems include an online service of the embedding model and a large-scale vector database. For traditional model upgrades, the old model will not be replaced by the new one until the embeddings of all the images in the database are re-computed by the new model, which takes days or weeks for a large amount of data. Recently, backward-compatible training (BCT) enables the new model to be immediately deployed online by making the new embeddings directly comparable to the old ones. For BCT, improving the compatibility of two models with less negative impact on retrieval performance is the key challenge. In this paper, we introduce AdvBCT, an Adversarial Backward-Compatible Training method with an elastic boundary constraint that takes both compatibility and discrimination into consideration. We first employ adversarial learning to minimize the distribution disparity between embeddings of the new model and the old model. Meanwhile, we add an elastic boundary constraint during training to improve compatibility and discrimination efficiently. Extensive experiments on GLDv2, Revisited Oxford (ROxford), and Revisited Paris (RParis) demonstrate that our method outperforms other BCT methods on both compatibility and discrimination. The implementation of AdvBCT will be publicly available at https://github.com/Ashespt/AdvBCT.
Tan Pan, Furong Xu, Sifeng He, Chen Jiang 0006, Qingpei Guo, Feng Qian 0006, Lei Yang 0061
CVPR1
2022 A Large-scale Comprehensive Dataset and Copy-overlap Aware Evaluation Protocol for Segment-level Video Copy Detection
abstract
In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level labelled data, with 160k realistic video copy pairs containing more than 280k localized copied segment pairs, but also covers a variety of video categories and a wide range of video duration. All the copied segments inside each collected video pair are manually extracted and accompanied by precisely annotated starting and ending timestamps. Alongside the dataset, we also propose a novel evaluation protocol that better measures the prediction accuracy of copy overlapping segments between a video pair and shows improved adaptability in different scenarios. By benchmarking several baseline and state-of-the-art segment-level video copy detection methods with the proposed dataset and evaluation metric, we provide a comprehensive analysis that uncovers the strengths and weaknesses of current approaches, hoping to open up promising directions for future works. The VCSL dataset, metric and benchmark codes are all publicly available at https://github.com/alipay/vCSL.
Sifeng He, Chen Jiang 0006, Gang Liang, Tan Pan, Qing Wang 0068, Furong Xu, Jingxiong Liu, Kaiming Huang, Feng Qian 0006, Lei Yang 0061
CVPR6
2021 Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
abstract
With the explosive growth of web videos in recent years, large-scale Content-Based Video Retrieval (CBVR) becomes increasingly essential in video filtering, recommendation, and copyright protection. Segment-level CBVR (S-CBVR) locates the start and end time of similar segments in finer granularity, which is beneficial for user browsing efficiency and infringement detection especially in long video scenarios. The challenge of S-CBVR task is how to achieve high temporal alignment accuracy with efficient computation and low storage consumption. In this paper, we propose a Segment Similarity and Alignment Network (SSAN) in dealing with the challenge which is firstly trained end-to-end in S-CBVR. SSAN is based on two newly proposed modules in video retrieval: (1) An efficient Self-supervised Keyframe Extraction (SKE) module to reduce redundant frame features, (2) A robust Similarity Pattern Detection (SPD) module for temporal alignment. In comparison with uniform frame extraction, SKE not only saves feature storage and search time, but also introduces comparable accuracy and limited extra computation time. In terms of temporal alignment, SPD localizes similar segments with higher accuracy and efficiency than existing deep learning methods. Furthermore, we jointly train SSAN with SKE and SPD and achieve an end-to-end improvement. Meanwhile, the two key modules SKE and SPD can also be effectively inserted into other video retrieval pipelines and gain considerable performance improvements. Experimental results on public datasets show that SSAN can obtain higher alignment accuracy while saving storage and online query computational cost compared to existing methods.
Chen Jiang 0006, Kaiming Huang, Sifeng He, Lei Yang 0061, Qing Wang 0068, Furong Xu, Tan Pan
ACM Multimedia11
2019 A Multi-Task Convolutional Neural Network for Renal Tumor Segmentation and Classification Using Multi-Phasic CT Images
abstract
Accounting for nearly 2% of all adults, renal cell carcinomas are sensitive to laparoscopic partial nephrectomy (LPN) which needs an accurate diagnosis and localization before operation. Faced with various intensity distribution, erratic location, irregular shape, etc, the image classification and semantic segmentation on CT scans of renal tumor are challenges. This paper presents a multi-task network, segmentation and classification convolutional neural network (SCNet), for preoperative assessment of renal tumor. Via the combination of two tasks, semantic features are fed to the classification network and classification results give segmentation network feedbacks in return. Besides, a 2-step segmentation strategy is conducted to the segmentation module which improves the result by 2.8%. Our experimental results of classification and segmentation achieve 100% accuracy and 0.882 dice coefficient of tumor region respectively, which are better than the results of a single classification network and segmentation network.
Tan Pan, Huazhong Shu, Jean-Louis Coatrieux, Guanyu Yang 0001, Chuanxia Wang, Ziwei Lu, Zhongwen Zhou, Youyong Kong, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger
ICIP1
2018 Automatic Segmentation of Kidney and Renal Tumor in CT Images Based on 3D Fully Convolutional Neural Network with Pyramid Pooling Module
abstract
Renal cancer is one of ten most common cancers in human beings. The laparoscopic partial nephrectomy (LPN) becomes the main therapeutic approach in treating renal cancer. Accurate kidney and tumor segmentation in CT images is a prerequisite step in the surgery planning. However, automatic and accurate kidney and renal tumor segmentation in CT images remains a challenge. In this paper, we propose a new method to perform a precise segmentation of kidney and renal tumor in CT angiography images. This method relies on a three-dimensional (3D) fully convolutional network (FCN) which combines a pyramid pooling module (PPM). The proposed network is implemented as an end-to-end learning system directly on 3D volumetric images. It can make use of the 3D spatial contextual information to improve the segmentation of the kidney as well as the tumor lesion. The experiments conducted on 140 patients show that these target structures can be segmented with a high accuracy. The resulting average dice coefficients obtained for kidney and renal tumor are equal to 0.931 and 0.802 respectively. These values are higher than those obtained from the other two neural networks.
Guanyu Yang 0001, Tan Pan, Youyong Kong, Jiasong Wu, Huazhong Shu, Limin Luo 0001, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Lijun Tang, Xiaomei Zhu
ICPR3