Yanpeng Sun

dblp:143/0055 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment Generation
abstract
Diffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. To address these limitations, we propose IMAGGarment+, an efficient diffusion framework for controllable and high-quality garment synthesis. It comprises two key modules designed for efficient and attribute-aware conditioning. First, we introduce an attribute-wise feature extractor (AFE) that disentangles key garment attributes, silhouette, logo, position, and color, into parallel latent streams. Each stream is optimized independently via LoRA, ensuring minimal parameter overhead while retaining expressive capacity. Second, we develop an attribute-adaptive attention (AA) module to inject attribute-specific cues into the generative process through a selective, layer-wise injection strategy. Specifically, silhouette and color features are injected into early decoder layers to guide structural and appearance formation, while logo features are propagated across all layers to ensure cross-scale consistency. Extensive experiments on fine-grained garment benchmarks demonstrate that IMAGGarment+ outperforms state-of-the-art baselines with less than 20% additional parameters, validating its effectiveness and efficiency.
Fei Shen 0004, Cong Wang 0034, Yanpeng Sun, Hao Tang 0007, Xiaoyu Du 0002
AAAI4
2026 GradAlign: Detecting Out-of-Distribution Samples via Gradient Concentration
Jiawei Gu, Yanpeng Sun, Hao Tang 0007, Zechao Li
Int. J. Comput. Vis.2
2026 CylindFormer: Image-to-Point Cloud Registration with Cylindrical Transformer
Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li
Int. J. Comput. Vis.3
2026 Refine, Control and Distill: A Text-to-Image Framework for Faithful Image Generation
abstract
While text-to-image diffusion models exhibit outstanding results, they struggle to faithfully generate key subjects with corresponding attributes in prompts, challenges known as catastrophic neglect and attribute binding. Previous works typically utilize attention adjustments to solve the above problems, whereas we observe that they may still generate unfaithful images. In this paper, we carefully analyze the text-to-image process and pinpoint three pivotal bottlenecks that hinder image faithful generation: (1) unequal responses of neglected subjects in text embedding, (2) competition and entanglement between subjects' attention, and (3) suboptimal quality of intermediate features from U-Net. Based on the aforementioned observations, we propose a Refine, Control, and Distill (RCD) framework built upon the stable diffusion model to alleviate the negative effects raised by the bottlenecks mentioned above, respectively. Specifically, we achieve the above goals through a text embedding refinement module, three region-level attention control losses, and self-distillation of intermediate semantic features in the denoising process. Our approach exhibits promising capability in generating faithful and high-quality images and outperforms state-of-the-art methods through extensive quantitative and qualitative evaluations on recent advanced base diffusion models.
Peng Xing, Ning Wang 0020, Yanpeng Sun, Jinhui Tang 0001, Zechao Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 FineG-RAG: Fine-Grained Retrieval-Augmented Generation for Multimodal Large Language Models
abstract
Fine-grained visual recognition refers to the ability to distinguish subtle differences between visually similar objects— a fundamental yet challenging capability for Multimodal Large Language Models (MLLMs). In this paper, we observe that even strong open-source MLLMs, such as Qwen2-VL and InternVL2, still struggle with accurately identifying fine-grained categories. These models often fail to attend to subtle but critical details for precise discrimination. To unlock this potential, we propose FineG-RAG, a retrieval-augmented generation pipeline designed to enhance the fine-grained recognition capabilities of MLLMs. FineG-RAG integrates external fine-grained knowledge into the recognition process via a generalized retriever. To support this, we construct fine-grained visual-language knowledge database containing representative images with wide visual diversity and expert-crafted attribute descriptions from multiple perspectives. Relevant fine-grained knowledge is retrieved from this database and fed into a visual-language augmented prompt, which provides rich multimodal context to guide MLLMs in generating accurate labels. To better evaluate the fine-grained recognition capabilities of MLLMs, we design a multiple-choice evaluation strategy based on publicly four fine-grained datasets. Extensive experiments demonstrate that FineG-RAG consistently outperforms baseline methods, achieving superior recognition accuracy across a range of off-the-shelf, open-source MLLMs.
Lu Jin 0001, Xinguang Xiang, Yanpeng Sun, Zechao Li, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 SSP-SAM: SAM With Semantic-Spatial Prompt for Referring Expression Segmentation
abstract
The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Expression Segmentation (RES). Toward this end, we propose SSP-SAM, a framework that fully utilizes SAM’s segmentation capabilities by integrating a Semantic-Spatial Prompt (SSP) encoder. Specifically, we incorporate both visual and linguistic attention adapters into the SSP encoder, which highlight salient objects within the visual features and discriminative phrases within the linguistic features. This design enhances the referent representation for the prompt generator, resulting in high-quality SSPs that enable SAM to generate precise masks guided by language. Although not specifically designed for Generalized RES (GRES), where the referent may correspond to zero, one, or multiple objects, SSP-SAM naturally supports this more flexible setting without additional modifications. Extensive experiments on widely used RES and GRES benchmarks confirm the superiority of our method. Notably, our approach generates segmentation masks of high quality, achieving strong precision even at strict thresholds such as [email protected]. Further evaluation on the PhraseCut dataset demonstrates improved performance in open-vocabulary scenarios compared to existing state-of-the-art RES methods. The code and checkpoints are available at: https://github.com/WayneTomas/SSP-SAM.
Wei Tang 0011, Xuejing Liu, Yanpeng Sun, Zechao Li
IEEE Trans. Circuits Syst. Video Technol.3
2026 Gradient Pruning Interactive Attack for Vision-Language Pre-Training Models
abstract
Vision-Language Pre-training (VLP) models exhibit pronounced vulnerability to multimodal adversarial examples, necessitating rigorous robustness research, particularly for transferable attacks in black-box scenarios. Current research predominantly enhances attack transferability across VLP models by diversifying image and text inputs. However, during adversarial example generation, these methods often prioritize amplifying inter-modal semantic discrepancies (i.e., modality-discrepancy features) while overlooking model-specific semantic features critical to transferable attacks. To address this limitation, we pro pose a transferable Gradient Pruning Interactive Attack (GPI Attack), which integrates gradient-pruned image perturbations with semantic-oriented text perturbations through modality interaction. For image attacks, extreme backpropagated gradients may cause adversarial examples to highlight certain model specific features, leading to poor transferability. To suppress this feature, the textual modality guides the pruning of extreme gradients within intermediate VLP blocks, and these pruned gradients are subsequently employed to direct the generation of adversarial images. For text attacks, we consolidate the perturbation process solely at the embedding level, which reduces semantic discrepancies across hierarchical structures and significantly enhances the generalizability of adversarial texts. Experimental results demonstrate the effectiveness of GPI-Attack in image text retrieval tasks on multimodal datasets such as Flickr30K and MSCOCO. Additionally, the proposed gradient pruning technique is plug-and-play, showing performance improvements even when applied to baseline methods, indicating its potential as a valuable enhancement for attack performance.
Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Zechao Li
IEEE Trans. Dependable Secur. Comput.3
2026 Visual Position Prompt for MLLM Based Visual Grounding
abstract
Although Multimodal Large Language Models (MLLMs) excel at various image-related tasks, they encounter challenges in precisely aligning coordinates with spatial information within images, particularly in position-aware tasks such as visual grounding. This limitation arises from two key factors. First, MLLMs lack explicit spatial references, making it difficult to associate textual descriptions with precise image locations. Second, their feature extraction processes prioritize global context over fine-grained spatial details, leading to weak localization capability. To address these issues, we introduce VPP-LLaVA, an MLLM enhanced with Visual Position Prompt (VPP) to improve its grounding capability. VPP-LLaVA integrates two complementary mechanisms: the global VPP overlays a learnable, axis-like tensor onto the input image to provide structured spatial cues, while the local VPP incorporates position-aware queries to support fine-grained localization. To effectively train our model with spatial guidance, we further introduce VPP-SFT, a curated dataset of 0.6 M high-quality visual grounding samples. Designed in a compact format, it enables efficient training and is significantly smaller than datasets used by other MLLMs (e.g., 21 M samples in MiniGPT-v2), yet still provides a strong performance boost. The resulting model, VPP-LLaVA, not only achieves state-of-the-art results on standard visual grounding benchmarks but also demonstrates strong zero-shot generalization to challenging unseen datasets.
Wei Tang 0011, Yanpeng Sun, Qinying Gu, Zechao Li
IEEE Trans. Multim.2
2025 Continual SFT Matches Multimodal RLHF with Negative Supervision
abstract
Multimodal RLHF usually happens after supervised fine-tuning (SFT) stage to continually improve vision-language models’ (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this preference alignment stage. In this paper, we observe that the inherent value of multimodal RLHF lies in its negative supervision, the logit of the rejected responses. We thus propose a novel negative supervised fine-tuning (nSFT) approach that fully excavates these information resided. Our nSFT disentangles this negative supervision in RLHF paradigm, and continually aligns VLMs with a simple SFT loss. This is more memory efficient than multimodal RLHF where 2 (e.g., DPO) or 4 (e.g., PPO) large VLMs are strictly required. The effectiveness of nSFT is rigorously proved by comparing it with various multimodal RLHF approaches, across different dataset sources, base VLMs and evaluation metrics. Besides, fruitful of ablations are provided to support our hypothesis. Code will be found in https://github.com/Kevinzcode/nSFT/.
Yanpeng Sun, Qiang Chen 0007, Jiangjiang Liu 0006, Jingdong Wang 0001
CVPR3
2025 DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback Synergy
abstract
Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the fundamental bottlenecks in existing RIS frameworks remains underexplored. To bridge this gap, we propose DeRIS, a novel framework that decomposes RIS into two key components: perception and cognition. This modular decomposition facilitates a systematic analysis of the primary bottlenecks impeding RIS performance. Our findings reveal that the predominant limitation lies not in perceptual deficiencies, but in the insufficient multi-modal cognitive capacity of current models. To mitigate this, we propose a Loopback Synergy mechanism, which enhances the synergy between the perception and cognition modules, thereby enabling precise segmentation while simultaneously improving robust image-text comprehension. Additionally, we analyze and introduce a simple non-referent sample conversion data augmentation to address the long-tail distribution issue related to target existence judgement in general scenarios. Notably, DeRIS demonstrates inherent adaptability to both non- and multi-referents scenarios without requiring specialized architectural modifications, enhancing its general applicability. The codes and models are available at https://github.com/Dmmm1997/DeRIS.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Wenxiao Cai, Yanpeng Sun, Wankou Yang
ICCV6
2025 Primitive Vision: Improving Diagram Understanding in MLLMs
abstract
Mathematical diagrams have a distinctive structure. Standard feature transforms designed for natural images (e.g., CLIP) fail to process them effectively, limiting their utility in multimodal large language models (MLLMs). Current efforts to improve MLLMs have primarily focused on scaling mathematical visual instruction datasets and strengthening LLM backbones, yet fine-grained visual recognition errors remain unaddressed. Our systematic evaluation on the visual grounding capabilities of state-of-the-art MLLMs highlights that fine-grained visual understanding remains a crucial bottleneck in visual mathematical reasoning (GPT-4o exhibits a 70% grounding error rate, and correcting these errors improves reasoning accuracy by 12%). We thus propose a novel approach featuring a geometrically-grounded vision encoder and a feature router that dynamically selects between hierarchical visual feature maps. Our model accurately recognizes visual primitives and generates precise visual prompts aligned with the language model’s reasoning needs. In experiments, PRIMITIVE-Qwen2.5-7B outperforms other 7B models by 12% on MathVerse and is on par with GPT-4V on MathVista. Our findings highlight the need for better fine-grained visual integration in MLLMs. Code is available at github.com/AI4Math-ShanZhang/SVE-Math.
Shan Zhang 0002, Aotian Chen, Yanpeng Sun, Jindong Gu, Yi-Yu Zheng, Piotr Koniusz, Anton van den Hengel, Yuan Xue 0002
ICML3
2025 FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
abstract
In this paper, we introduce FedMGP, a new paradigm for personalized federated prompt learning in vision-language models (VLMs). Existing federated prompt learning (FPL) methods often rely on a single, text-only prompt representation, which leads to client-specific overfitting and unstable aggregation under heterogeneous data distributions. Toward this end, FedMGP equips each client with multiple groups of paired textual and visual prompts, enabling the model to capture diverse, fine-grained semantic and instance-level cues. A diversity loss is introduced to drive each prompt group to specialize in distinct and complementary semantic aspects, ensuring that the groups collectively cover a broader range of local characteristics.During communication, FedMGP employs a dynamic prompt aggregation strategy based on similarity-guided probabilistic sampling: each client computes the cosine similarity between its prompt groups and the global prompts from the previous round, then samples s groups via a softmax-weighted distribution. This soft selection mechanism preferentially aggregates semantically aligned knowledge while still enabling exploration of underrepresented patterns—effectively balancing the preservation of common knowledge with client-specific features. Notably, FedMGP maintains parameter efficiency by redistributing a fixed prompt capacity across multiple groups, achieving state-of-the-art performance with the lowest communication parameters (5.1k) among all federated prompt learning methods. Theoretical analysis shows that our dynamic aggregation strategy promotes robust global representation learning by reinforcing shared semantics while suppressing client-specific noise. Extensive experiments demonstrate that FedMGP consistently outperforms prior approaches in both personalization and domain generalization across diverse federated vision-language benchmarks.The code will be released on https://github.com/weihao-bo/FedMGP.git.
Weihao Bo, Yanpeng Sun, Xinyu Zhang 0017, Zechao Li
NeurIPS2
2025 Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
abstract
Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular, panoramic X-rays, a widely used imaging modality in oral radiology, pose interpretative challenges due to dense anatomical structures and subtle pathological cues, which are not captured by existing medical benchmarks or instruction datasets. To this end, we introduce MMOral, the first large-scale multimodal instruction dataset and benchmark tailored for panoramic X-ray interpretation. MMOral consists of 20,563 annotated images paired with 1.3 million instruction-following instances across diverse task types, including attribute extraction, report generation, visual question answering, and image-grounded dialogue. In addition, we present MMOral-Bench, a comprehensive evaluation suite covering five key diagnostic dimensions in dentistry. We evaluate 64 LVLMs on MMOral-Bench and find that even the best-performing model, i.e., GPT-4o, only achieves 43.31% accuracy, revealing significant limitations of current models in this domain. To promote the progress of this specific domain, we provide the supervised fine-tuning (SFT) process utilizing our meticulously curated MMOral instruction dataset. Remarkably, a single epoch of SFT yields substantial performance enhancements for LVLMs, e.g., Qwen2.5-VL-7B demonstrates a 24.73% improvement. MMOral holds significant potential as a critical foundation for intelligent dentistry and enables more clinically impactful multimodal AI systems in the dental field.
Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Qi Yong H. Ai, Lun M. Wong, Hao Tang 0005, Kuo Feng Hung
NeurIPS3
2025 CSGO: Content-Style Composition in Text-to-Image Generation
abstract
The advancement of image style transfer has been fundamentally constrained by the absence of large-scale, high-quality datasets with explicit content-style-stylized supervision. Existing methods predominantly adopt training-free paradigms (e.g., image inversion), which limit controllability and generalization due to the lack of structured triplet data. To bridge this gap, we design a scalable and automated pipeline that constructs and purifies high-fidelity content-style-stylized image triplets. Leveraging this pipeline, we introduce IMAGStyle—the first large-scale dataset of its kind, containing 210K diverse and precisely aligned triplets for style transfer research. Empowered by IMAGStyle, we propose CSGO, a unified, end-to-end trainable framework that decouples content and style representations via independent feature injection. CSGO jointly supports image-driven style transfer, text-driven stylized generation, and text-editing-driven stylized synthesis within a single architecture. Extensive experiments show that CSGO achieves state-of-the-art controllability and fidelity, demonstrating the critical role of structured synthetic data in unlocking robust and generalizable style transfer. Source code: \url{https://github.com/instantX-research/CSGO}
Peng Xing, Yanpeng Sun, Qixun Wang 0001, Hao Ai, Jen-Yuan Huang, Zechao Li
NeurIPS3
2025 SSA: semantic structure aware inference on CNN networks for weakly pixel-wise dense predictions without cost
Yanpeng Sun, Zechao Li
Frontiers Comput. Sci.1
2025 Modality-Specific Interactive Attack for Vision-Language Pre-Training Models
abstract
Recent advances have heightened the interest in the adversarial transferability of Vision-Language Pre-training (VLP) models. However, most existing strategies constrained by two persistent limitations: suboptimal utilization of cross-modal interactive information, and inherent discrepancies across hierarchical textual representation. To address these challenges, we propose the Modality-Specific Interactive Attack (MSI-Attack), a novel approach that integrates semantic-level image perturbations with embedding-level text perturbations, all while maintaining minimal inter-modal constraints. In our image attack methodology, we introduce Multi-modal Integrated Gradients (MIG) to guide perturbations toward the core semantics of images, enriched by their associated deeply text information. This technique enhances transferability by capturing consistent features across various models, thereby effectively misleading similar-model perception areas. Additionally, we employ a momentum iteration strategy in conjunction with MIG, which amalgamates current and historical gradients to expedite the perturbation updates. For text attacks, we streamline the perturbation process by operating exclusively at the embedding level. This reduces semantic gaps across hierarchical structures and significantly enhances the generalizability of adversarial text. Moreover, we delve deeper into how semantic perturbations with varying degrees of similarity affect the overall attack effectiveness. Our experimental results on image-text retrieval tasks using the multi-modal datasets Flickr30K and MSCOCO underscore the efficacy of MSI-Attack. Our method achieves superior performance, setting a new state-of-the-art benchmark, all without the need for additional mechanisms.
Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li
IEEE Trans. Inf. Forensics Secur.3
2025 Exploring Effective Factors for Improving Visual In-Context Learning
abstract
The In-Context Learning (ICL) is to understand a new task via a few demonstrations (aka. prompt) and predict new inputs without tuning the models. While it has been widely studied in NLP, it is still a relatively new area of research in computer vision. To reveal the factors influencing the performance of visual in-context learning, this paper shows that Prompt Selection and Prompt Fusion are two major factors that have a direct impact on the inference performance of visual in-context learning. Prompt selection is the process of selecting the most suitable prompt for query image. This is crucial because high-quality prompts assist large-scale visual models in rapidly and accurately comprehending new tasks. Prompt fusion involves combining prompts and query images to activate knowledge within large-scale visual models. However, altering the prompt fusion method significantly impacts its performance on new tasks. Based on these findings, we propose a simple framework prompt-SelF to improve visual in-context learning. Specifically, we first use the pixel-level retrieval method to select a suitable prompt, and then use different prompt fusion methods to activate diverse knowledge stored in the large-scale vision model, and finally, ensemble the prediction results obtained from different prompt fusion methods to obtain the final prediction results. We conducted extensive experiments on single-object segmentation and detection tasks to demonstrate the effectiveness of prompt-SelF. Remarkably, prompt-SelF has outperformed OSLSM method-based meta-learning in 1-shot segmentation for the first time. This indicated the great potential of visual in-context learning. The source code and models will be available at https://github.com/syp2ysy/prompt-SelF.
Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Jingdong Wang 0001, Zechao Li
IEEE Trans. Image Process.1
2024 VRP-SAM: SAM with Visual Reference Prompt
abstract
In this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Any-thing Model (SAM) to utilize annotated reference images as prompts for segmentation, creating the VRP-SAM model. In essence, VRP-SAM can utilize annotated reference images to comprehend specific objects and perform segmen-tation of specific objects in target image. It is note that the VRP encoder can support a variety of annotation for-mats for reference images, including point, box, scribble, and mask. VRP-SAM achieves a breakthrough within the SAM framework by extending its versatility and applicabil-ity while preserving SAM's inherent strengths, thus enhancing user-friendliness. To enhance the generalization abil-ity of VRP-SAM, the VRP encoder adopts a meta-learning strategy. To validate the effectiveness of VRP-SAM, we con-ducted extensive empirical studies on the Pascal and COCO datasets. Remarkably, VRP-SAM achieved state-of-the-art performance in visual reference segmentation with mini-mal learnable parameters. Furthermore, VRP-SAM demon-strates strong generalization capabilities, allowing it to per-form segmentation of unseen objects and enabling cross-domain segmentation. The source code and models will be available at https://github.com/syp2ysy/VRP-SAM
Yanpeng Sun, Shan Zhang 0002, Xinyu Zhang 0017, Qiang Chen 0007, Errui Ding, Jingdong Wang 0001, Zechao Li
CVPR1
2024 Clutter Suppression for Through-the-Wall Radar Based on Robust Non-Negative Matrix Factorization in Dual-Domain
abstract
The strong clutter typically impedes the accurate imaging and detection of targets for through-the-wall (TWR) system. In this letter, a joint dual-domain (JDD) clutter suppression method based on robust non-negative matrix factorization (RNMF) is proposed for TWR system. The proposed method exploits RNMF algorithm to remove the clutter in both the signal and image domains. Specifically, the exponentially weighted multiplication fusion is employed to fuse the two decluttered images in the signal and image domains. The optimal value of the exponential factor for fusion is determined using the minimum entropy criterion. The experimental results have shown that the proposed method can provide the better clutter suppression performance compared to the existing low rank and sparse decomposition (LRSD) based approaches.
Lele Qu, Qiyue Hu, Tianhong Yang, Yanpeng Sun
IEEE Geosci. Remote. Sens. Lett.5
2024 Normal Image Guided Segmentation Framework for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection is required to detect/segment anomalous samples/regions that deviate from the normal pattern while learning only through the normal sample category. Towards this end, this paper proposes a novel framework for anomaly detection by introducing normal images as guidance called Normal Image Guided Segmentation Framework (NIGSF). It consists of a Normal Guided Network (NGN) and a Saliency Augmentation Module (SAM). NGN constructs the contrast set, which is a candidate set for extracting normal sample features. Then, a normal feature extractor is developed to extract detailed and complete features containing normal semantic information as guidance features. Meanwhile, the guidance feature fusion module is introduced to realize normal semantic guidance in the feature space, and then the segmentation module discriminates the features that are different from the normal guidance features as anomalies. SAM aims to generate forged anomaly samples utilizing available normal samples. It introduces saliency maps and random Perlin noise to generate saliency Perlin noise maps and then to generate diverse forged anomaly samples. Extensive experiments are conducted to evaluate the performance of NIGSF on three anomaly detection benchmark datasets. The results demonstrate the effectiveness of each proposed module and the superiority of the proposed method. Specifically, NIGSF outperforms the runner-up by 5.4% in terms of anomaly segmentation AP metric.
Peng Xing, Yanpeng Sun, Dan Zeng 0001, Zechao Li
IEEE Trans. Circuits Syst. Video Technol.2
2023 σ-Adaptive Decoupled Prototype for Few-Shot Object Detection
abstract
Meta-learning-based few-shot detectors use one K-average-pooled prototype (averaging along K-shot dimension) in both Region Proposal Network (RPN) and Detection head (DH) for query detection. Such plain operation would harm the FSOD performance in two aspects: 1) the poor quality of the prototype, and 2) the equivocal guidance due to the contradictions between RPN and DH. In this paper, we look closely into those critical issues and propose the σ-Adaptive Decoupled Prototype (σ-ADP) as a solution. To generate the high-quality prototype, we prioritize salient representations and deemphasize trivial variations by accessing both angle distance and magnitude dispersion (σ) across K-support samples. To provide precise information for the query image, the prototype is decoupled into task-specific ones, which provide tailored guidance for ‘where to look’ and ‘what to look for’, respectively.Beyond that, we find our σ-ADP can gradually strengthen the generalization power of encoding network during meta-training. So it can robustly deal with intra-class variations and a simple K- average pooling is enough to generate a high-quality prototype at meta-testing. We provide theoretical analysis to support its rationality. Extensive experiments on Pascal VOC, MS-COCO and FSOD datasets demonstrate that the proposed method achieves new state-of-the-art performance. Notably, our method surpasses the baseline model by a large margin – up to around 5.0% AP50and 8.0% AP75on novel classes.
Jinhao Du, Shan Zhang 0002, Qiang Chen 0007, Haifeng Le, Yanpeng Sun, Yao Ni, Jian Wang 0066, Jingdong Wang 0001
ICCV5
2023 Who is partner: A new perspective on data association of multi-object tracking
Yuqing Ding, Yanpeng Sun, Zechao Li
Image Vis. Comput.2
2022 Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuning
abstract
Freezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, leading to better model generalization on learning novel classes. Our method decomposes backbone parameters into three successive matrices via the Singular Value Decomposition (SVD), then {\em only fine-tunes the singular values} and keeps others frozen. The above design allows the model to adjust feature representations on novel classes while maintaining semantic clues within the pre-trained backbone. We evaluate our {\em Singular Value Fine-tuning (SVF)} approach on various few-shot segmentation methods with different backbones. We achieve state-of-the-art results on both Pascal-5$^i$ and COCO-20$^i$ across 1-shot and 5-shot settings. Hopefully, this simple baseline will encourage researchers to rethink the role of backbone fine-tuning in few-shot settings.
Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng 0001, Zechao Li, Jingdong Wang 0001
NeurIPS1
2022 Considering Fine-Grained and Coarse-Grained Information for Context-Aware Recommendations
abstract
Abstract In context-aware recommendation systems, most existing methods encode users’ preferences by mapping item and category information into the same space, which is just a stack of information. The item and category information contained in the interaction behaviours is not fully utilized. Moreover, since users’ preferences for a candidate item are influenced by the changes in temporal and historical behaviours, it is unreasonable to predict correlations between users and candidates by using users’ fixed features. A fine-grained and coarse-grained information based framework proposed in our paper which considers multi-granularity information of users’ historical behaviours. First, a parallel structure is provided to mine users’ preference information under different granularities. Then, self-attention and attention mechanisms are used to capture the dynamic preferences. Experiment results on two publicly available datasets show that our framework outperforms state-of-the-art methods across the calculated evaluation metrics.
Yiqin Luo, Yanpeng Sun, Liang Chang 0003, Tianlong Gu, Chenzhong Bin, Long Li 0005
Comput. J.2
2022 Enhanced Through-the-Wall Radar Imaging Based on Deep Layer Aggregation
abstract
The accurate imaging of stationary human targets in the indoor scene containing strong scatterers such as cabinets, tables, and chairs is very important for the through-the-wall radar (TWR) system. The convolution neural network (CNN) has been used to enhance radar imaging quality. In this letter, a novel multiresolution fusion network (MRFN) based on the deep layer aggregation (DLA) method is proposed for TWR imaging. The proposed MRFN can accurately localize the weak scattering human targets and provide the scattering intensity differences between the strong and weak targets. Both the simulated and real TWR data are used to evaluate the imaging performance of the proposed MRFN. The experimental results demonstrate the superiority of the proposed TWR imaging method over the existing CNN-based imaging methods.
Lele Qu, Changan Wang, Tianhong Yang, Lili Zhang 0005, Yanpeng Sun
IEEE Geosci. Remote. Sens. Lett.5
2022 CTNet: Context-Based Tandem Network for Semantic Segmentation
abstract
Contextual information has been shown to be powerful for semantic segmentation. This work proposes a novel Context-based Tandem Network (CTNet) by interactively exploring the spatial contextual information and the channel contextual information, which can discover the semantic context for semantic segmentation. Specifically, the Spatial Contextual Module (SCM) is leveraged to uncover the spatial contextual dependency between pixels by exploring the correlation between pixels and categories. Meanwhile, the Channel Contextual Module (CCM) is introduced to learn the semantic features including the semantic feature maps and class-specific features by modeling the long-term semantic dependence between channels. The learned semantic features are utilized as the prior knowledge to guide the learning of SCM, which can make SCM obtain more accurate long-range spatial dependency. Finally, to further improve the performance of the learned representations for semantic segmentation, the results of the two context modules are adaptively integrated to achieve better results. Extensive experiments are conducted on four widely-used datasets, i.e., PASCAL-Context, Cityscapes, ADE20K and PASCAL VOC2012. The results demonstrate the superior performance of the proposed CTNet by comparison with several state-of-the-art methods. The source code and models are available at https://github.com/syp2ysy/CTNet.
Zechao Li, Yanpeng Sun, Liyan Zhang 0001, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 WGAN-GP-Based Synthetic Radar Spectrogram Augmentation in Human Activity Recognition
abstract
Despite deep convolutional neural networks (DCNNs) having been used extensively in radar-based human activity recognition in recent years, their performance could not be fully implemented because of the lack of radar dataset. However, radar data acquisition is difficult to achieve due to the high cost of its measurement. Generative adversarial networks (GANs) can be utilized to generate a large number of similar micro-Doppler signatures with which to increase the training data set. For the training of DCNNs, the quality and diversity of data set generated by GANs is particularly important. In this paper, we propose using a more stable and effective Wasserstein generative adversarial network with gradient penalty (WGAN-GP) to augment the training data set. The classification results from the experimental data have shown the proposed method can improve the classification accuracy of human activity.
Lele Qu, Tianhong Yang, Lili Zhang 0005, Yanpeng Sun
IGARSS5
2019 Sparse Recovery Method for Estimation of Wall Parameters in Through-the-Wall Radar
abstract
Estimating unknown wall parameters is of great importance for the application of through-the-wall radar (TWR). The time-delay-only estimation (TDOE) method is able to efficiently retrieve the constitute parameters of the homogeneous wall under test. For the TDOE method, the estimation accuracy of time delays associated with wall reflections directly affects the accuracy of wall parameters estimation. In this paper, we propose a sparse recovery method for estimation of wall parameters, which utilizes the orthogonal matching pursuit (OMP) algorithm to perform the time delay estimation of echoes backscattered from the wall at each antenna separation. Numerical simulation results have shown that the proposed estimation method is capable of providing the estimation of wall parameters with higher accuracy.
Lele Qu, Zhongli Fang, Tianhong Yang, Yanpeng Sun, Lili Zhang 0005
IGARSS4
2019 A Neural User Preference Modeling Framework for Recommendation Based on Knowledge Graph
Guiming Zhu, Chenzhong Bin, Tianlong Gu, Liang Chang 0003, Yanpeng Sun, Wei Chen 0105, Zhonghao Jia
PRICAI (1)5
2019 Through-the-wall radar imaging algorithm for moving target under wall parameter uncertainties
abstract
In order to solve the problems of slow imaging speed and poor reconstruction accuracy of wall parameters under the condition of wall parameter fuzziness, an improved limited Broyden–Fletcher–Goldfarb–Shanno‐particle swarm optimisation (LBFGS‐PSO) algorithm was proposed. The LBFGS‐PSO algorithm model solves the problems of slow calculation speed and large errors of the traditional quasi‐Newton algorithm and particle swarm algorithm. The algorithm combined with block orthogonal matching pursuit algorithm can not only accurately reconstruct the position of the sidewall, but also can use the multi‐path information to accurately reconstruct the moving target and the stationary target. Compared with the traditional BFGS algorithm and PSO algorithm, the proposed algorithm can reduce the calculation time and provide more accurate estimation results. Simulation results and data analysis verify the performance of the proposed algorithm.
Yanpeng Sun, Lele Qu
IET Image Process.1
2019 A personalized POI route recommendation system based on heterogeneous tourism data and sequential pattern mining
Chenzhong Bin, Tianlong Gu, Yanpeng Sun, Liang Chang 0003
Multim. Tools Appl.3
2019 A Travel Route Recommendation System Based on Smart Phones and IoT Environment
abstract
Tourism recommendation systems play a vital role in providing useful travel information to tourists. However, existing systems rarely aim at recommending tangible itineraries for tourists within a specific POI due to their lack of onsite travel behavioral data and related route mining algorithms. To this end, a novel travel route recommendation system is proposed, which collects tourist onsite travel behavior data automatically regarding a specific POI based on smart phone and IoT technology. Then, the proposed system preprocesses the behavior data to transform raw behavior sequences into Tourist-Behavior pattern sequences. Subsequently, the system discovers frequent travel routes from the generated pattern sequences by using an original route mining algorithm, named Tourist-Behavior PrefixSpan. Finally, a route-recommending method is designed to search and rank tangible travel routes according to the querying tourist’s profile and constraint. The experimental results demonstrate that the proposed system is efficient and effective in recommending POI-oriented tangible travel routes considering tourists’ route constraints and personal profile while ensuring that the suggested routes have considerable route values.
Chenzhong Bin, Tianlong Gu, Yanpeng Sun, Liang Chang 0003
Wirel. Commun. Mob. Comput.3
2018 Personalized POIs Travel Route Recommendation System Based on Tourism Big Data
Chenzhong Bin, Tianlong Gu, Yanpeng Sun, Liang Chang 0003, Wenping Sun
PRICAI3
2018 A Multi-latent Semantics Representation Model for Mining Tourist Trajectory
Yanpeng Sun, Tianlong Gu, Chenzhong Bin, Liang Chang 0003, Haili Kuang, Zhaowei Huang
PRICAI (1)1
2016 MT-BCS-Based Two-Dimensional Diffraction Tomographic GPR Imaging Algorithm With Multiview-Multistatic Configuration
abstract
High-resolution ground-penetrating radar multiview-multistatic diffraction-tomographic (DT) imaging usually requires the wide signal bandwidth and large antenna aperture, which results in the great amount of imaging data. To solve the aforementioned problem, an innovative 2-D multiview-multistatic DT imaging algorithm based on the multitask Bayesian compressive sensing (MT-BCS) strategy is proposed in this letter. The reduction of the measurement data can be achieved by performing a reduced set of measurements in the frequency domain. In particular, a joint Bayesian sparse reconstruction scheme is used to recover the original frequency domain data from the reduced frequency measurements across all the measurement positions. Finally, the image of the investigation domain can be reconstructed by the traditional multiview-multistatic DT imaging algorithm. Numerical simulation results have shown that the proposed imaging method can not only reduce the frequency measurement data but also provide the satisfactory quality of the reconstructed image.
Yanpeng Sun, Lele Qu, Yuqing Yin
IEEE Geosci. Remote. Sens. Lett.1
2015 Diffraction Tomographic Ground-Penetrating Radar Multibistatic Imaging Algorithm With Compressive Frequency Measurements
abstract
High-resolution diffraction tomographic (DT) ground-penetrating radar (GPR) image formation requires the use of wideband signal and large antenna array aperture, which leads to the generation of large amounts of imaging data. A compressive sensing multibistatic GPR DT imaging algorithm is presented in this letter. The proposed imaging algorithm can provide the advantage in terms of reducing the measured data in the frequency domain while maintaining the image quality of the reconstructed scenario. The imaging results reconstructed via the processing of synthetic data have verified the validity and effectiveness of the proposed imaging method.
Lele Qu, Yuqing Yin, Yanpeng Sun, Lili Zhang 0005
IEEE Geosci. Remote. Sens. Lett.3
2014 Time-Delay Estimation for Ground Penetrating Radar Using ESPRIT With Improved Spatial Smoothing Technique
abstract
Estimating the time delays of buried target echoes is particularly important for the application of ground penetrating radar (GPR). Due to its smaller computational burden, the estimation of signal parameters via rotational invariance technique (ESPRIT) is preferred to process the buried target echoes in the frequency domain and to obtain the accurate super-resolution time delays. In this letter, we give an in-depth analysis of the essential preprocessing steps for the application of ESPRIT to practical GPR measurement data. In particular, an improved spatial smoothing method is adopted to construct the correlation matrix for the robustness of the time-delay estimation result. The effectiveness of the algorithm is verified by the synthetic data from a horizontally stratified medium model using the finite-difference time-domain method, which explicitly takes into account surface scattering for a more realistic scenario.
Lele Qu, Tianhong Yang, Lili Zhang 0005, Yanpeng Sun
IEEE Geosci. Remote. Sens. Lett.5