VLDB 2026 Research / reviewers in the wild / expert
Jinpeng Li 0004
dblp:95/2448-4
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-1752-8883ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language ModelabstractVision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialities, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters. Yaoqian Li, Xikai Yang, Dunyuan Xu, Litao Zhao, Xiaowei Hu 0001, Jinpeng Li 0004, Pheng-Ann Heng |
AAAI | 7 |
| 2025 | Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
Dunyuan Xu, Yuchen Yuan, Xikai Yang, Jingyang Zhang, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (16) | 6 |
| 2025 | Medical Large Vision Language Models with Multi-image Visual Ability
Xikai Yang, Juzheng Miao, Yuchen Yuan, Qi Dou 0001, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (5) | 6 |
| 2025 | Gated-GPS: enhancing protein-protein interaction site prediction with scalable learning and imbalance-aware optimizationabstractIn protein-protein interaction site (PPIS) prediction, existing machine learning models struggle with small datasets, limiting their predictive accuracy for unseen proteins. Additionally, class imbalance in protein complexes, where binding residues constitute a small fraction of all residues, hinders model performance. To address these challenges, we constructed a training dataset 9$\times $ larger than previous benchmarks by filtering the latest protein-protein complex data, improving diversity and generalization. We propose Gated-GPS, a Graph Transformer model with a novel gating mechanism designed to effectively leverage this expanded dataset. Additionally, we integrate cross-entropy loss with Tversky Loss to adjust sensitivity to positive and negative samples, mitigating class imbalance by emphasizing underrepresented binding residues. Experimental results show that Gated-GPS outperforms state-of-the-art (SOTA) models across four test sets. Notably, on the UBTest dataset, designed to evaluate generalization on unbounded proteins, our method improves MCC and AUPRC by 18.5% and 21.4%, respectively, over the previous SOTA. In a case study of snake venom toxin-protein interactions, our model accurately identified interaction sites, demonstrating its potential for therapeutic design and advancing the understanding of complex protein interactions. Xin Gao 0026, Hanqun Cao, Jinpeng Li 0004, Jiezhong Qiu, Guangyong Chen, Pheng-Ann Heng |
Briefings Bioinform. | 3 |
| 2025 | Multi-Scale Spatio-Temporal Transformer-Based Imbalanced Longitudinal Learning for Glaucoma Forecasting From Irregular Time Series ImagesabstractGlaucoma is one of the major eye diseases that leads to progressive optic nerve fiber damage and irreversible blindness, afflicting millions of individuals. Glaucoma forecast is a good solution to early screening and intervention of potential patients, which is helpful to prevent further deterioration of the disease. It leverages a series of historical fundus images of an eye and forecasts the likelihood of glaucoma occurrence in the future. However, the irregular sampling nature and the imbalanced class distribution are two challenges in the development of disease forecasting approaches. To this end, we introduce the Multi-scale Spatio-temporal Transformer Network (MST-former) based on the transformer architecture tailored for sequential image inputs, which can effectively learn representative semantic information from sequential images on both temporal and spatial dimensions. Specifically, we employ a multi-scale structure to extract features at various resolutions, which can largely exploit rich spatial information encoded in each image. Besides, we design a time distance matrix to scale time attention in a non-linear manner, which could effectively deal with the irregularly sampled data. Furthermore, we introduce a temperature-controlled Balanced Softmax Cross-entropy loss to address the class imbalance issue. Extensive experiments on the Sequential fundus Images for Glaucoma Forecast (SIGF) dataset demonstrate the superiority of the proposed MST-former method, achieving an AUC of 96.6% for glaucoma forecasting. Besides, our method shows excellent generalization capability on the Alzheimer's Disease Neuroimaging Initiative (ADNI) MRI dataset, with an accuracy of 88.2% for mild cognitive impairment and Alzheimer's disease prediction, outperforming the compared method by a large margin. A series of ablation studies further verify the contribution of our proposed components in addressing the irregular sampled and class imbalanced problems. Xikai Yang, Xi Wang 0013, Yuchen Yuan, Jinpeng Li 0004, Guangyong Chen, Ning Li Wang, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | PAL: Boosting Skin Lesion Segmentation via Probabilistic Attribute LearningabstractSkin lesion segmentation is vital for the early detection, diagnosis, and treatment of melanoma, yet it remains challenging due to significant variations in lesion attributes (e.g., color, size, shape), ambiguous boundaries, and noise interference. Recent advancements have focused on capturing contextual information and incorporating boundary priors to handle challenging lesions. However, there has been limited exploration on the explicit analysis of the inherent patterns of skin lesions, a crucial aspect of the knowledge-driven decision-making process used by clinical experts. In this work, we introduce a novel approach called Probabilistic Attribute Learning (PAL), which leverages knowledge of lesion patterns to achieve enhanced performance on challenging lesions. Recognizing that the lesion patterns exhibited in each image can be properly depicted by disentangled attributes, we begin by explicitly estimating the distributions of these attributes as distinct Gaussian distributions, with mean and variance indicating the most likely pattern of that attribute and its variation. Using Monte Carlo Sampling, we iteratively draw multiple samples from these distributions to capture various potential patterns for each attribute. These samples are then merged through an effective attribute fusion technique, resulting in diverse representations that comprehensively depict the lesion class. By performing pixel-class proximity matching between each pixel-wise representation and the diverse class-wise representations, we significantly enhance the model's robustness. Extensive experiments on two public skin lesion datasets and one unified polyp lesion dataset demonstrate the effectiveness and strong generalization ability of our method. Codes are available at https://github.com/IsYuchenYuan/PAL. Yuchen Yuan, Xi Wang 0013, Jinpeng Li 0004, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2024 | PointPatchMix: Point Cloud Mixing with Patch ScoringabstractData augmentation is an effective regularization strategy for mitigating overfitting in deep neural networks, and it plays a crucial role in 3D vision tasks, where the point cloud data is relatively limited. While mixing-based augmentation has shown promise for point clouds, previous methods mix point clouds either on block level or point level, which has constrained their ability to strike a balance between generating diverse training samples and preserving the local characteristics of point clouds. The significance of each part component of the point clouds has not been fully considered, as not all parts contribute equally to the classification task, and some parts may contain unimportant or redundant information. To overcome these challenges, we propose PointPatchMix, a novel approach that mixes point clouds at the patch level and integrates a patch scoring module to generate content-based targets for mixed point clouds. Our approach preserves local features at the patch level, while the patch scoring module assigns targets based on the content-based significance score from a pre-trained teacher model. We evaluate PointPatchMix on two benchmark datasets including ModelNet40 and ScanObjectNN, and demonstrate significant improvements over various baselines in both synthetic and real-world datasets, as well as few-shot settings. With Point-MAE as our baseline, our model surpasses previous methods by a significant margin. Furthermore, our approach shows strong generalization across various point cloud methods and enhances the robustness of the baseline model. Code is available at https://jiazewang.com/projects/pointpatchmix.html. Jinpeng Li 0004, Guangyong Chen, Anfeng Liu, Pheng-Ann Heng |
AAAI | 3 |
| 2024 | SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
Hao Chen 0193, Jinpeng Li 0004, Chenyong Guan, Guangyong Chen, Pheng-Ann Heng |
BMVC | 4 |
| 2024 | Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
Jinpeng Li 0004, Jiancheng Huang, Qiang Nie, Yong Liu 0032, Bin-Bin Gao, Qiong Wang 0001, Pheng-Ann Heng, Guangyong Chen |
BMVC | 3 |
| 2024 | 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
Shizhan Gong, Yuan Zhong 0003, Wenao Ma, Jinpeng Li 0004, Zhao Wang 0006, Jingyang Zhang, Pheng-Ann Heng, Qi Dou 0001 |
Medical Image Anal. | 4 |
| 2023 | Fast Non-Markovian Diffusion Model for Weakly Supervised Anomaly Detection in Brain MR Images
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (5) | 1 |
| 2023 | Learning Robust Classifier for Imbalanced Medical Image Dataset with Noisy Labels by Minimizing Invariant Risk
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (6) | 1 |
| 2023 | Efficient Person Search: An Anchor-Free Approach
Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Peng Zheng 0004, Shengcai Liao, Xiaokang Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2022 | Exploring Visual Context for Weakly Supervised Person SearchabstractPerson search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive, limiting the practicability and scalability of current frameworks. This paper inventively considers weakly supervised person search with only bounding box annotations. We propose to address this novel task by investigating three levels of context clues (i.e., detection, memory and scene) in unconstrained natural images. The first two are employed to promote local and global discriminative capabilities, while the latter enhances clustering accuracy. Despite its simple design, our CGPS boosts the baseline model by 8.8% in mAP on CUHK-SYSU. Surprisingly, it even achieves comparable performance with several supervised person search models. Our code is available at https://github. com/ljpadam/CGPS. Yichao Yan, Jinpeng Li 0004, Shengcai Liao, Jie Qin 0004, Bingbing Ni, Ke Lu 0002, Xiaokang Yang 0001 |
AAAI | 2 |
| 2022 | RePFormer: Refinement Pyramid Transformer for Robust Facial Landmark DetectionabstractThis paper presents a Refinement Pyramid Transformer (RePFormer) for robust facial landmark detection. Most facial landmark detectors focus on learning representative image features. However, these CNN-based feature representations are not robust enough to handle complex real-world scenarios due to ignoring the internal structure of landmarks, as well as the relations between landmarks and context. In this work, we formulate the facial landmark detection task as refining landmark queries along pyramid memories. Specifically, a pyramid transformer head (PTH) is introduced to build both homologous relations among landmarks and heterologous relations between landmarks and cross-scale contexts. Besides, a dynamic landmark refinement (DLR) module is designed to decompose the landmark regression into an end-to-end refinement procedure, where the dynamically aggregated queries are transformed to residual coordinates predictions. Extensive experimental results on four facial landmark detection benchmarks and their various subsets demonstrate the superior performance and high robustness of our framework. Jinpeng Li 0004, Haibo Jin, Shengcai Liao, Ling Shao 0001, Pheng-Ann Heng |
IJCAI | 1 |
| 2022 | Flat-Aware Cross-Stage Distilled Framework for Imbalanced Medical Image Classification
Jinpeng Li 0004, Guangyong Chen, Hangyu Mao, Danruo Deng, Dong Li 0016, Jianye Hao, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 1 |
| 2022 | Urban scene based Semantical Modulation for Pedestrian Detection
Hangzhi Jiang, Shengcai Liao, Jinpeng Li 0004, Véronique Prinet, Shiming Xiang |
Neurocomputing | 3 |
| 2021 | Generalizable Pedestrian Detection: The Elephant in the RoomabstractPedestrian detection is used in many vision based applications ranging from video surveillance to autonomous driving. Despite achieving high performance, it is still largely unknown how well existing detectors generalize to unseen data. This is important because a practical detector should be ready to use in various scenarios in applications. To this end, we conduct a comprehensive study in this paper, using a general principle of direct cross-dataset evaluation. Through this study, we find that existing state-of-the-art pedestrian detectors, though perform quite well when trained and tested on the same dataset, generalize poorly in cross dataset evaluation. We demonstrate that there are two reasons for this trend. Firstly, their designs (e.g. anchor settings) may be biased towards popular benchmarks in the traditional single-dataset training and test pipeline, but as a result largely limit their generalization capability. Secondly, the training source is generally not dense in pedestrians and diverse in scenarios. Under direct cross-dataset evaluation, surprisingly, we find that a general purpose object detector, without pedestrian-tailored adaptation in design, generalizes much better compared to existing state-of-the-art pedestrian detectors. Furthermore, we illustrate that diverse and dense datasets, collected by crawling the web, serve to be an efficient source of pre-training for pedestrian detection. Accordingly, we propose a progressive training pipeline and find that it works well for autonomous-driving oriented pedestrian detection. Consequently, the study conducted in this paper suggests that more emphasis should be put on cross-dataset evaluation for the future design of generalizable pedestrian detectors. Code and models can be accessed at https://github.com/hasanirtiza/Pedestron. Irtiza Hasan, Shengcai Liao, Jinpeng Li 0004, Saad Ullah Akram, Ling Shao 0001 |
CVPR | 3 |
| 2021 | Anchor-Free Person SearchabstractPerson search aims to simultaneously localize and identify a query person from realistic, uncropped images, which can be regarded as the unified task of pedestrian detection and person re-identification (re-id). Most existing works employ two-stage detectors like Faster-RCNN, yielding encouraging accuracy but with high computational overhead. In this work, we present the Feature-Aligned Person Search Network (AlignPS), the first anchor-free framework to efficiently tackle this challenging task. AlignPS explicitly addresses the major challenges, which we summarize as the misalignment issues in different levels (i.e., scale, region, and task), when accommodating an anchor-free detector for this task. More specifically, we propose an aligned feature aggregation module to generate more discriminative and robust feature embeddings by following a "re-id first" principle. Such a simple design directly improves the baseline anchor-free model on CUHK-SYSU by more than 20% in mAP. Moreover, AlignPS outperforms state-of-the-art two-stage methods, with a higher speed. The code is available at https://github.com/daodaofr/AlignPS. Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Song Bai 0001, Shengcai Liao, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
CVPR | 2 |
| 2020 | Box Guided Convolution for Pedestrian DetectionabstractOcclusions, scale variation and numerous false positives still represent fundamental challenges in pedestrian detection. Intuitively, different sizes of receptive fields and more attention to the visible parts are required for detecting pedestrians with various scales and occlusion levels, respectively. However, these challenges have not been addressed well by existing pedestrian detectors. This paper presents a novel convolutional network, denoted as box guided convolution network (BGCNet), to tackle these challenges simultaneously in a unified framework. In particular, we proposed a box guided convolution (BGC) that can dynamically adjust the sizes of convolution kernels guided by the predicted bounding boxes. In this way, BGCNet provides position-aware receptive fields to address the challenge of large variations of scales. In addition, for the issue of heavy occlusion, the kernel parameters of BGC are spatially localized around the salient and mostly visible key points of a pedestrian, such as the head and foot, to effectively capture high-level semantic features to help detection. Furthermore, a local maximum (LM) loss is introduced to depress false positives and highlight true positives by forcing positives, rather than negatives, as local maximums, without any additional inference burden. We evaluate BGCNet on popular pedestrian detection benchmarks, and achieve the state-of-the-art results, with the significant performance improvement on heavily occluded and small-scale pedestrians. Jinpeng Li 0004, Shengcai Liao, Hangzhi Jiang, Ling Shao 0001 |
ACM Multimedia | 1 |