VLDB 2026 Research / reviewers in the wild / expert
Jinpeng Chen 0003
dblp:91/10208-3
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-0469-4463ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding boxes to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap raises uncertainty about whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and utilize them to solve problems. To address this limitation, we introduce VP-Bench, aiming to assess MLLMs’ capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models’ ability to perceive VPs in natural scenes, utilizing 100K visualized prompts spanning 8 shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 21 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL-2.5 and Qwen2.5-VL). In addition, we conduct a comprehensive analysis of the factors influencing VP understanding, such as attribute variations and model scale. VP-Bench establishes a new reference framework for studying MLLMs’ ability to comprehend and resolve grounded referring questions. Mingjie Xu, Jinpeng Chen 0003, Yuzhi Zhao, Jason Chun Lok Li, Zekang Du, Mengyang Wu, Kun Li 0015, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei |
AAAI | 2 |
| 2025 | KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented GenerationabstractZiyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jason Chun Lok Li, Zhijian Hou, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen 0003, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong 0001 |
EMNLP | 8 |
| 2025 | SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction TuningabstractMultimodal Continual Instruction Tuning (MCIT) aims to enable Multimodal Large Language Models (MLLMs) to incrementally learn new tasks without catastrophic forgetting, thus adapting to evolving requirements. In this paper, we explore the forgetting caused by such incremental training, categorizing it into superficial forgetting and essential forgetting. Superficial forgetting refers to cases where the model’s knowledge may not be genuinely lost, but its responses to previous tasks deviate from expected formats due to the influence of subsequent tasks’ answer styles, making the results unusable. On the other hand, essential forgetting refers to situations where the model provides correctly formatted but factually inaccurate answers, indicating a true loss of knowledge. Assessing essential forgetting necessitates addressing superficial forgetting first, as severe superficial forgetting can conceal the model’s knowledge state. Hence, we first introduce the Answer Style Diversification (ASD) paradigm, which defines a standardized process for data style transformations across different tasks, unifying their training sets into similarly diversified styles to prevent superficial forgetting caused by style shifts. Building on this, we propose RegLoRA to mitigate essential forgetting. RegLoRA stabilizes key parameters where prior knowledge is primarily stored by applying regularization to LoRA’s weight update matrices, enabling the model to retain existing competencies while remaining adaptable to new tasks. Experimental results demonstrate that our overall method, SEFE, achieves state-of-the-art performance. Jinpeng Chen 0003, Runmin Cong, Yuzhi Zhao, Hongzheng Yang, Guang-Neng Hu, Horace Ho-Shing Ip, Sam Kwong |
ICML | 1 |
| 2025 | MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationabstractDespite the effectiveness of Segment Anything Model (SAM) based methods in Few-Shot Segmentation (FSS) tasks, our closer examination of their prompt encoding mechanism reveals that these methods rely solely on visual information to generate a single type of prompt. Consequently, they suffer from semantic granularity representation bias and a loss of spatial information. To address these limitations, this paper introduces an innovative multi-modal prompt encoder, enabling SAM to leverage both annotated reference images and textual descriptions of class names as segmentation prompts. This approach generates text prompts, dense visual prompts, and sparse visual prompts, spanning multiple modalities and granularities. These prompts provide enhanced representations of the target class, capturing both abstract semantics and specific details, while ensuring granularity appropriateness. When our multi-modal prompt encoder is integrated with SAM's image encoder and mask decoder, the overall model is referred to as MM-Prompt. To validate its effectiveness, we conducted extensive empirical studies on the PASCAL-5^i and COCO-20^i datasets. The experimental results demonstrate that MM-Prompt achieves state-of-the-art performance in FSS tasks, highlighting its substantial potential and value in this domain. Runmin Cong, Jinpeng Chen 0003, Chen Zhang 0013, Feng Li 0037, Huihui Bai 0001, Sam Kwong |
ACM Multimedia | 3 |
| 2025 | SLRanger: an integrated approach for spliced leader detection and operon prediction using long RNA readsabstractSpliced leader (SL) trans-splicing occurs in a wide range of eukaryotes and plays a critical role in processing mRNAs derived from operon structures. However, current research on this mechanism remains limited, partly due to the difficulty in accurately identifying genuine SL trans-splicing events. The advent of long-read RNA sequencing technologies, such as direct RNA sequencing by Oxford Nanopore Technologies, offers a more promising avenue for detecting these events with greater resolution. Here, we present SLRanger, an integrated tool to detect SL sequences and predict operon structures in eukaryotic transcriptomes. SLRanger improves upon the traditional Smith-Waterman (SW) alignment framework by incorporating an optimized scoring scheme tailored to SL detection in native long RNA reads. We primarily validated our method using direct RNA sequencing data from Caenorhabditis elegans, a well-established model organism for studying trans-splicing. Through a dynamic cutoff strategy, SLRanger robustly identified high-confidence SL-carrying reads. Leveraging the SL information, SLRanger achieved over 80% accuracy in operon gene prediction, recovering more than 70% of known operon genes in C. elegans. SLRanger was also applied to detect SL from cDNA long RNA reads and another trans-spliced species. Our results demonstrate that SLRanger not only provides a reliable approach for characterizing SL trans-splicing events but also serves as an effective framework for operon discovery, enabling transcriptomic analysis for operons and facilitating downstream data-mining applications. Yanwen Shao, Jinpeng Chen 0003 |
Briefings Bioinform. | 3 |
| 2025 | Replay Without Saving: Prototype Derivation and Distribution Rebalance for Class-Incremental Semantic SegmentationabstractThe research of class-incremental semantic segmentation (CISS) seeks to enhance semantic segmentation methods by enabling the progressive learning of new classes while preserving knowledge of previously learned ones. A significant yet often neglected challenge in this domain is class imbalance. In CISS, each task focuses on different foreground classes, with the training set for each task exclusively comprising images that contain these currently focused classes. This results in an overrepresentation of these classes within the single-task training set, leading to a classification bias towards them. To address this issue, we propose a novel CISS method named STAR, whose core principle is to reintegrate the missing proportions of previous classes into current single-task training samples by replaying their prototypes. Moreover, we develop a prototype deviation technique that enables the deduction of past-class prototypes, integrating the recognition patterns of the classifiers and the extraction patterns of the feature extractor. With this technique, replay can be accomplished without using any storage to save prototypes. Complementing our method, we devise two loss functions to enforce cross-task feature constraints: the Old-Class Features Maintaining (OCFM) loss and the Similarity-Aware Discriminative (SAD) loss. The OCFM loss is designed to stabilize the feature space of old classes, thus preserving previously acquired knowledge without compromising the ability to learn new classes. The SAD loss aims to enhance feature distinctions between similar old and new class pairs, minimizing potential confusion. Our experiments on two public datasets, Pascal VOC 2012 and ADE20 K, demonstrate that our STAR achieves state-of-the-art performance. Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Trace Back and Go Ahead: Completing partial annotation for continual semantic segmentationabstractExisting Continual Semantic Segmentation (CSS) methods effectively address the issue of background shift in regular training samples. However, this issue persists in exemplars, i.e. , replay samples, which is often overlooked. Each exemplar is annotated only with the classes from its originating task, while other past classes and the current classes during replay are labeled as background . This partial annotation can erase the network’s knowledge of previous classes and impede the learning of new classes. To resolve this, we introduce a new method named Trace Back and Go Ahead (TAGA), which utilizes a backward annotator model and a forward annotator model to generate pseudo-labels for both regular training samples and exemplars, aiming at reducing the adverse effects of incomplete annotations. This approach effectively mitigates the risk of incorrect guidance from both sample types, offering a comprehensive solution to background shift . Additionally, due to a significantly smaller number of exemplars compared to regular training samples, the class distribution in the sample pool of each incremental task exhibits a long-tailed pattern, potentially biasing classification towards incremental classes. Consequently, TAGA incorporates a class-equilibrium sampling strategy that adaptively adjusts the sampling frequencies based on the ratios of exemplars to regular samples and past to new classes, counteracting the skewed distribution. Extensive experiments on two public datasets, Pascal VOC 2012 and ADE20K, demonstrate that our method surpasses state-of-the-art methods. • Proposes a method to address background shift problems in all training samples. • Utilizes annotators to complete missing annotations for past and new classes. • Implements a class-equilibrium sampling strategy to long-tail challenges. • Demonstrates superior performance of TAGA over the state-of-the-art CSS methods. Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
Pattern Recognit. | 2 |
| 2025 | Concept-Level Semantic Transfer and Context-Level Distribution Modeling for Few-Shot SegmentationabstractFew-shot segmentation (FSS) methods aim to segment objects using only a few pixel-level annotated samples. Current approaches either derive a generalized class representation from support samples to guide the segmentation of query samples, which often discards crucial spatial contextual information, or rely heavily on spatial affinity between support and query samples, without adequately summarizing and utilizing the core information of the target class. Consequently, the former struggles with fine detail accuracy, while the latter tends to produce errors in overall localization. To address these issues, we propose a novel FSS framework, CCFormer, which balances the transmission of core semantic concepts with the modeling of spatial context, improving both macro and micro-level segmentation accuracy. Our approach introduces three key modules: 1) the Concept Perception Generation (CPG) module, which leverages pre-trained category perception capabilities to capture high-quality core representations of the target class; 2) the Concept-Feature Integration (CFI) module, which injects the core class information into both support and query features during feature extraction; and 3) the Contextual Distribution Mining (CDM) module, which utilizes a Brownian Distance Covariance matrix to model the spatial-channel distribution between support and query samples, preserving the fine-grained integrity of the target. Experimental results on the PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that CCFormer achieves state-of-the-art performance, with visualizations further validating its effectiveness. Our code is available at github.com/lourise/ccformer. Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Strike a Balance in Continual Panoptic Segmentation
Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
ECCV (41) | 1 |
| 2024 | KepSalinst: Using Peripheral Points to Delineate Salient InstancesabstractSalient instance segmentation (SIS) is an emerging field that evolves from salient object detection (SOD), aiming at identifying individual salient instances using segmentation maps. Inspired by the success of dynamic convolutions in segmentation tasks, this article introduces a keypoints-based SIS network (KepSalinst). It employs multiple keypoints, that is, the center and several peripheral points of an instance, as effective geometrical guidance for dynamic convolutions. The features at peripheral points can help roughly delineate the spatial extent of the instance and complement the information inside the central features. To fully exploit the complementary components within these features, we design a differentiated patterns fusion (DPF) module. This ensures that the resulting dynamic convolutional filters formed by these features are sufficiently comprehensive for precise segmentation. Furthermore, we introduce a high-level semantic guided saliency (HSGS) module. This module enhances the perception of saliency by predicting a map for the input image to estimate a saliency score for each segmented instance. On four SIS datasets (ILSO, SOC, SIS10K, and COME15K), our KepSalinst outperforms all previous models qualitatively and quantitatively. Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
IEEE Trans. Cybern. | 1 |
| 2024 | Query-Guided Prototype Evolution Network for Few-Shot SegmentationabstractPrevious Few-Shot Segmentation (FSS) approaches exclusively utilize support features for prototype generation, neglecting the specific requirements of the query. To address this, we present the Query-guided Prototype Evolution Network (QPENet), a new method that integrates query features into the generation process of foreground and background prototypes, thereby yielding customized prototypes attuned to specific queries. The evolution of the foreground prototype is accomplished through a support-query-support iterative process involving two new modules: Pseudo-prototype Generation (PPG) and Dual Prototype Evolution (DPE). The PPG module employs support features to create an initial prototype for the preliminary segmentation of the query image, resulting in a pseudo-prototype reflecting the unique needs of the current query. Subsequently, the DPE module performs reverse segmentation on support images using this pseudo-prototype, leading to the generation of evolved prototypes, which can be considered as custom solutions. As for the background prototype, the evolution begins with a global background prototype that represents the generalized features of all training images. We also design a Global Background Cleansing (GBC) module to eliminate potential adverse components mirroring the characteristics of the current foreground class. Experimental results on the PASCAL-52and COCO-202datasets attest to the substantial enhancements achieved by QPENet over prevailing state-of-the-art techniques, underscoring the validity of our ideas. Runmin Cong, Jinpeng Chen 0003, Wei Zhang 0021, Qingming Huang, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow DetectionabstractDespite significant progress in shadow detection, current methods still struggle with the adverse impact of background color, which may lead to errors when shadows are present on complex backgrounds. Drawing inspiration from the human visual system, we treat the input shadow image as a composition of a background layer and a shadow layer, and design a Style-guided Dual-layer Disentanglement Network (SDDNet) to model these layers independently. To achieve this, we devise a Feature Separation and Recombination (FSR) module that decomposes multi-level features into shadow-related and background-related components by offering specialized supervision for each component, while preserving information integrity and avoiding redundancy through the reconstruction constraint. Moreover, we propose a Shadow Style Filter (SSF) module to guide the feature disentanglement by focusing on style differentiation and uniformization. With these two modules and our overall pipeline, our model effectively minimizes the detrimental effects of background color, yielding superior performance on three public datasets with a real-time inference speed of 32 FPS. Our code is publicly available at:https://github.com/rmcong/SDDNet_ACMMM23. Runmin Cong, Yuchen Guan, Jinpeng Chen 0003, Wei Zhang 0021, Yao Zhao 0001, Sam Kwong |
ACM Multimedia | 3 |
| 2023 | Saving 100x Storage: Prototype Replay for Reconstructing Training Sample Distribution in Class-Incremental Semantic SegmentationabstractExisting class-incremental semantic segmentation (CISS) methods mainly tackle catastrophic forgetting and background shift, but often overlook another crucial issue. In CISS, each step focuses on different foreground classes, and the training set for a single step only includes images containing pixels of the current foreground classes, excluding images without them. This leads to an overrepresentation of these foreground classes in the single-step training set, causing the classification biased towards these classes. To address this issue, we present STAR, which preserves the main characteristics of each past class by storing a compact prototype and necessary statistical data, and aligns the class distribution of single-step training samples with the complete dataset by replaying these prototypes and repeating background pixels with appropriate frequency. Compared to the previous works that replay raw images, our method saves over 100 times the storage while achieving better performance. Moreover, STAR incorporates an old-class features maintaining (OCFM) loss, keeping old-class features unchanged while preserving sufficient plasticity for learning new classes. Furthermore, a similarity-aware discriminative (SAD) loss is employed to specifically enhance the feature diversity between similar old-new class pairs. Experiments on two public datasets, Pascal VOC 2012 and ADE20K, reveal that our model surpasses all previous state-of-the-art methods. Jinpeng Chen 0003, Runmin Cong, Horace Ho-Shing Ip, Sam Kwong |
NeurIPS | 1 |