Shanmin Pang

dblp:129/3978 · DBLP profile ↗
← Back
59ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0001-7217-864XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems
abstract
Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scenarios. Existing patch-based attack methods are primarily designed for object detection models. Due to the more complex architectures and strong reasoning capabilities of MLLMs, these approaches perform poorly when transferred to MLLM-based systems. To address these limitations, we propose PhysPatch, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems. PhysPatch jointly optimizes patch location, shape, and content to enhance attack effectiveness and real-world applicability. It introduces a semantic-based mask initialization strategy for realistic placement, an SVD-based local alignment loss with patch-guided crop-resize to improve transferability, and a potential field-based mask refinement method. Extensive experiments across open-source, commercial, and reasoning-capable MLLMs demonstrate that PhysPatch significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs. Moreover, PhysPatch consistently places adversarial patches in physically feasible regions of AD scenes, ensuring strong real-world applicability and deployability.
Qi Guo 0008, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang 0026, Ju Jia, Yang Liu 0003, Qing Guo 0005
AAAI3
2026 Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and Prunable
abstract
Query-based models are extensively used in 3D object detection tasks, with a wide range of pre-trained checkpoints readily available online. However, despite their popularity, these models often require an excessive number of object queries, far surpassing the actual number of objects to detect. The redundant queries result in unnecessary computational and memory costs. In this paper, we find that not all queries contribute equally -- a significant portion of queries have a much smaller impact compared to others. Based on this observation, we propose an embarrassingly simple approach called Gradually Pruning Queries (GPQ), which prunes queries incrementally based on their classification scores. A key advantage of GPQ is that it requires no additional learnable parameters. It is straightforward to implement in any query-based method, as it can be seamlessly integrated as a fine-tuning step using an existing checkpoint after training. With GPQ, users can easily generate multiple models with fewer queries, starting from a checkpoint with an excessive number of queries. Experiments on various advanced 3D detectors show that GPQ effectively reduces redundant queries while maintaining performance. Using our method, model inference on desktop GPUs can be accelerated by up to 1.35x. Moreover, after deployment on edge devices, it achieves up to a 67.86% reduction in FLOPs and a 65.16% decrease in inference time.
Lizhen Xu, Wenzhao Qiu, Shanmin Pang, Xiuxiu Bai, Jianru Xue
AAAI4
2026 Deep Reinforcement Learning-Based Experience Sharing for On-Ramp Merging in Mixed Traffic Environments
Fuhao Liu, Shenye Dong, Shanmin Pang
IV5
2026 Structure searchable network model for deepfake image detection
Zinian Liu, Ningning Bai, Ruidong Han, Shanmin Pang
Pattern Recognit.6
2026 Image splicing localization method driven by device difference feature guidance
Ningning Bai, Ruidong Han, Jianpeng Hou, Tongtong Xu, Shanmin Pang
Pattern Recognit.8
2026 Unsupervised Domain Adaptation-Based Cross-Type Deepfake Image Detection
abstract
In practical applications of social media and the Internet, deepfake face images involve a plethora of unlabeled samples. To effectively identify unlabeled deepfake images, the domain adaptation technique has gained significant attention. It applies the knowledge learned from labeled samples (source domain) to unlabeled samples (target domain) in a cross-domain manner. However, the existing domain adaptation-based deepfake detection methods primarily focus on intra-type cross-domain scenarios. In this study, we propose an unsupervised domain adaptation-based deepfake face image detection method for extra-type cross-domain scenarios. The core idea of our approach lies in the development of a domain adaptation model that consists of Domain Tag Adversarial (DTA) and Domain Feature Alignment (DFA) algorithms, called DTA-DFA, which empowers the proposed method with strong cross-domain capability. The DTA is utilized to weaken the specificity within each domain, while DFA aligns the distribution between the source and target domains. Compared with the existing deepfake detection methods, the experimental results demonstrate that the proposed method dramatically enhances the extra-type cross-domain detection performance. Moreover, the DTA-DFA model also exhibits a remarkable ability to perform cross-domain detection from large-shot labeled samples to few-shot labeled samples, further verifying its powerful cross-domain capability. Code is released at https://github.com/QinQin741/DTA-DFA-DA-model.
Zinian Liu, Ningning Bai, Minghua Zhao, Shanmin Pang
IEEE Trans. Image Process.6
2026 A Human-Oriented Cooperative Driving Approach: Integrating Driving Intention, State, and Conflict
abstract
Human-vehicle cooperative driving serves as a vital bridge to fully autonomous driving by improving driving flexibility and gradually building driver trust and acceptance of autonomous technology. To establish more natural and effective human-vehicle interaction, we propose a Human-Oriented Cooperative Driving (HOCD) approach that primarily minimizes human-machine conflict by prioritizing driver intention and state. In implementation, we take both tactical and operational levels into account to ensure seamless human-vehicle cooperation. At the tactical level, we design an intention-aware trajectory planning method, using intention consistency cost as the core metric to evaluate the trajectory and align it with driver intention. At the operational level, we develop a control authority allocation strategy based on reinforcement learning, optimizing the policy through a designed reward function to achieve consistency between driver state and authority allocation. The results of simulation and human-in-the-loop experiments demonstrate that our proposed approach not only aligns with driver intention in trajectory planning but also ensures a reasonable authority allocation. Compared to other cooperative driving approaches, the proposed HOCD approach significantly enhances driving performance and mitigates human-machine conflict.
Shanmin Pang, Jianwu Fang, Shengye Dong, Fuhao Liu, Jianru Xue, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.2
2026 MDT-FI: Mask-Guided Dual-Branch Transformer With Texture and Structure Feature Interaction for Image Inpainting
abstract
Image inpainting has attracted considerable attention in computer vision and image processing due to its wide range of applications. While deep learning-based methods have shown promising potential, accurately recovering pixel-level details remains a significant challenge, particularly in the presence of large and irregular missing regions. Furthermore, existing methods are limited by unidirectional semantic guidance and a localized understanding of global structural context. In this study, we propose a mask-guided dual-branch Transformer-based framework, named MDT-FI, which effectively balances local detail restoration and global contextual reasoning by explicitly modeling long-range dependencies. MDT-FI consists of three key components: the Interactive Attention Module (IAM), the Spectral Harmonization Module (SHM), and the Lateral Adaptation Network (LAN). The model integrates multi-scale feature interaction, frequency-domain information fusion, and a mask-guided attention mechanism to progressively build cross-level feature associations. This design facilitates multi-level representation learning and optimization, thereby enhancing local texture synthesis while preserving global structural consistency. To further improve perceptual quality, a feature augmenter is employed to assess the fidelity of both texture and structure in the generated results. Extensive experiments on CelebA-HQ, Places2, and Paris Street View demonstrate that MDT-FI significantly outperforms state-of-the-art methods.
Dong Liu 0043, Ruidong Han, Jianghua Li, Shanmin Pang
IEEE Trans. Multim.5
2025 Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua
ICCV4
2025 Accelerate 3D Object Detection Models via Zero-Shot Attention Key Pruning
abstract
Query-based methods with dense features have demonstrated remarkable success in 3D object detection tasks. However, the computational demands of these models, particularly with large image sizes and multiple transformer layers, pose significant challenges for efficient running on edge devices. Existing pruning and distillation methods either need retraining or are designed for ViT models, which are hard to migrate to 3D detectors. To address this issue, we propose a zero-shot runtime pruning method for transformer decoders in 3D object detection models. The method, termed tgGBC (trim keys gradually Guided By Classification scores), systematically trims keys in transformer modules based on their importance. We expand the classification score to multiply it with the attention map to get the importance score of each key and then prune certain keys after each transformer layer according to their importance scores. Our method achieves a 1.99x speedup in the transformer decoder of the latest ToC3D model, with only a minimal performance loss of less than 1%. Interestingly, for certain models, our method even enhances their performance. Moreover, we deploy 3D detectors with tgGBC on an edge device, further validating the effectiveness of our method. The code can be found at https://github.com/iseri27/tg_gbc.
Lizhen Xu, Xiuxiu Bai, Xiaojun Jia, Jianwu Fang, Shanmin Pang
ICCV5
2025 HeightMapNet: Explicit Height Modeling for End-to-End HD Map Learning
Wenzhao Qiu, Shanmin Pang, Jianwu Fang, Jianru Xue
WACV2
2025 Towards robust DeepFake distortion attack via adversarial autoaugment
Qi Guo 0008, Shanmin Pang, Qing Guo 0005
Neurocomputing2
2025 Towards generalizable face forgery detection via mitigating spurious correlation
Ningning Bai, Ruidong Han, Jianpeng Hou, Shanmin Pang
Neural Networks6
2025 PIM-Net: Progressive Inconsistency Mining Network for image manipulation localization
Ningning Bai, Ruidong Han, Jianpeng Hou, Shanmin Pang
Pattern Recognit.6
2025 Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models via Diffusion Models
abstract
Adversarial attacks, particularly targeted transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential security flaws before deployment. However, previous transfer-based adversarial attacks incur high costs due to high iteration counts and complex method structure. Furthermore, due to the unnaturalness of adversarial semantics, the generated adversarial examples have low transferability. These issues limit the utility of existing methods for assessing robustness. To address these issues, we propose AdvDiffVLM, which uses diffusion models to generate natural, unrestricted and targeted adversarial examples via score matching. Specifically, AdvDiffVLM uses Adaptive Ensemble Gradient Estimation (AEGE) to modify the score during the diffusion model’s reverse generation process, ensuring that the produced adversarial examples have natural adversarial targeted semantics, which improves their transferability. Simultaneously, to improve the quality of adversarial examples, we use the GradCAM-guided Mask Generation (GCMG) to disperse adversarial semantics throughout the image rather than concentrating them in a single area. Finally, AdvDiffVLM embeds more target semantics into adversarial examples after multiple iterations. Experimental results show that our method generates adversarial examples 5x to 10x faster than state-of-the-art (SOTA) transfer-based adversarial attacks while maintaining higher quality adversarial examples. Furthermore, compared to previous transfer-based adversarial attacks, the adversarial examples generated by our method have better transferability. Notably, AdvDiffVLM can successfully attack a variety of commercial VLMs in a black-box environment, including GPT-4V. The code is available athttps://github.com/gq-max/AdvDiffVLM
Qi Guo 0008, Shanmin Pang, Xiaojun Jia, Yang Liu 0003, Qing Guo 0005
IEEE Trans. Inf. Forensics Secur.2
2024 ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep Generation
abstract
The commercial text-to-image deep generation models (e.g. DALL·E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent jailbreaking methods generate adversarial prompts capable of bypassing safety filters and producing unsafe content, exposing vulnerabilities in influential commercial models. However, once these adversarial prompts are identified, the safety filter can be updated to prevent the generation of unsafe images. In this work, we propose an effective, simple, and difficult-to-detect jailbreaking solution: generating safe content initially with normal text prompts and then editing the generations to embed unsafe content. The intuition behind this idea is that the deep generation model cannot reject safe generation with normal text prompts, while the editing models focus on modifying the local regions of images and do not involve a safety strategy. However, implementing such a solution is non-trivial, and we need to overcome several challenges: how to automatically confirm the normal prompt to replace the unsafe prompts, and how to effectively perform editable replacement and naturally generate unsafe content. In this work, we propose the collaborative generation and editing for jailbreaking text-to-image deep generation (ColJailBreak), which comprises three key components: adaptive normal safe substitution, inpainting-driven injection of unsafe content, and contrastive language-image-guided collaborative optimization. We validate our method on three datasets and compare it to two baseline methods. Our method could generate unsafe content through two commercial deep generation models including GPT-4 and DALL·E 2.
Yizhuo Ma, Shanmin Pang, Qi Guo 0008, Qing Guo 0005
NeurIPS2
2024 Image splicing region localization with adaptive multi-feature filtration
Jianpeng Hou, Ruidong Han, Mao Jia, Dong Liu 0043, Qinhua Yu, Shanmin Pang
Expert Syst. Appl.7
2024 Physically Driven Self-Supervised Learning and its Applications in Geophysical Inversion
abstract
Sparse coding (SC) has been proven effective in various geological tasks, such as seismic time-frequency (TF) analysis and seismic reflection inversion. Nevertheless, it inevitably has several drawbacks, e.g., low computational efficiency and difficulty in parameter selection. Recently, self-supervised learning (SSL) has emerged as a promising alternative to mitigate these issues, offering high computational effectiveness and requiring fewer labels. We suggest a generalized physically driven workflow for geophysical inversion based on SSL and SC, named the physically driven SSL network (PDSSLNet). This generalized PDSSLNet model comprises two main modules. One is the inverse model, generated by convolutional neural networks (CNNs), which can benefit from their high computational effectiveness and strong nonlinear fitting ability. The other one is the forward model based on the SC theory, ensuring the physical meaning of the geophysical applications with high accuracy. Afterward, we provide two typical geological inversion cases to demonstrate the validity and effectiveness of the suggested PDSSLNet, including sparse TF analysis and seismic reflectivity inversion. Three-dimensional (3D) field data volume applications confirm that the proposed inversion workflow may efficiently circumvent the drawbacks of the conventional SC-based approach while maintaining excellent computing efficiency.
Yang Yang 0069, Naihao Liu, Shanmin Pang, Rongchang Liu, Jinghuai Gao
IEEE Trans. Geosci. Remote. Sens.5
2024 CTE-Net: Contextual Texture Enhancement Network for Image Super-Resolution
abstract
The object of image super-resolution reconstruction is to overcome the limitations imposed by hardware imaging conditions and patterns, aiming to restore high-frequency details in images through signal processing techniques. Recently, deep learning-based single-image super-resolution reconstruction (SISR) has achieved remarkable performance. However, the current methods exhibit inadequate performance in the reconstruction of texture details, thereby posing a challenge for further enhancing the accuracy of super-resolution reconstruction. In this study, we propose a novel contextual texture enhancement network (CTE-Net) aimed at improving the level of texture details in image super-resolution. The CTE-Net comprises of two crucial components: the multi-level feature aggregation module (MFAM) and the contextual information enhancement module (CIEM). The MFAM integrates global and local low-resolution (LR) features from both the pixel space and channel dimensions, thereby enhancing the feature representation capability of the network. The CIEM is deployed to enhance the network's learning capacity by integrating a meticulously designed context-attention mechanism, which effectively explores the adjacent contextual information of images and thereby amplifies the expressive capability of the generated features. Moreover, we utilize local binary patterns (LBP) to guide the feature selection strategies for MFAM and CIEM, thereby prioritizing the network's decision logic towards the recovery of texture details. The extensive experiments demonstrate that our method yields satisfactory results. In comparison to the state-of-the-art approaches, our method exhibits superior performance on the benchmark datasets.
Dong Liu 0043, Ruidong Han, Ningning Bai, Jianpeng Hou, Shanmin Pang
IEEE Trans. Multim.6
2024 A Mutually Textual and Visual Refinement Network for Image-Text Matching
abstract
Image-text matching is vital important in the field of multi-modal intelligence. Recently, it is advocated in a way that decomposes images and texts into local fragments and followed by region-word aligning. As a result, the image-text relevance score is given by aggregating semantic similarities between matched region-word pairs. Despite effectiveness, this strategy fails to express data relations exactly. From the perspective of the text side, text words decomposed from a concise language sentence usually have limited contextual information, which can result in semantic identical but actually false text-region alignments. From the perspective of the image side, semantic ambiguity that multiple objects share the same semantic meaning can further exacerbate this problem. In this manuscript, we introduce a mutually Textual and Visual Refinement Network (TVRN), to tackle the inaccurate cross-modal alignment problem. In a nutshell, TVRN improves inter-modal matching by improving contextual information in sentences meanwhile reduces semantic ambiguity in images to capture the maximized relevant relations. More specifically, we develop a new module that integrates visual contextual clues into the text modality to generate informational text features with richer geometric contexts. Mutually, we further design a semantic alignment enhancement module that leverages consensus affinity of local image and text features to guide deeper semantic image embedding with the supervision of global image vectors. At the image-text matching stage, similarities at the local and global levels are integrated to capture coarse-grained and fine-grained interactions between vision and language. A large number of experiments on Flickr30K and MS-COCO benchmarks demonstrate that TVRN is superior to existing methods.
Shanmin Pang, Yueyang Zeng, Jianru Xue
IEEE Trans. Multim.1
2023 Feature fine-tuning and attribute representation transformation for zero-shot learning
Shanmin Pang, Wenyu Hao, Yang Long 0001
Comput. Vis. Image Underst.1
2023 Hierarchical block aggregation network for long-tailed visual recognition
Shanmin Pang, Wei Wang 0016, Renzhong Zhang, Wenyu Hao
Neurocomputing1
2023 Tensor-Based Incomplete Multi-View Clustering With Low-Rank Data Reconstruction and Consistency Guidance
abstract
We propose a new approach, called Tensor-based Incomplete Multi-view Clustering with Low-rank data Reconstruction and Consistency guidance (TIMC-RC), to perform clustering on multi-view data with missing views. Existing methods usually leverage original incomplete data to explore the partial correlations among multiple views, and do not make sufficient use of both consistent and complementary information across views. To explore the full information of missing and available views, TIMC-RC introduces low-rank data reconstruction and consistency view establishment. Specifically, 1) it adopts a low-rank constraint to reconstruct data representations so as to reduce the negative effect of missing data and obtain more reasonable data representations. 2) It builds a new consistency view by self-representation matrices and therefore explores the consistent correlation of different views. 3) It formalizes view-specific self-representation matrices and the consistent matrix as a tensor and utilizes the tensor singular value decomposition-based nuclear norm to enhance the consistency and complementarity of multi-view representations. Experiments conducted on eight benchmarks verify the effectiveness and advancement of the proposed TIMC-RC.
Wenyu Hao, Shanmin Pang, Xiuxiu Bai, Jianru Xue
IEEE Trans. Circuits Syst. Video Technol.2
2022 Tensor-based multi-view clustering with consistency exploration and diversity regularization
Wenyu Hao, Shanmin Pang, Bo Yang 0041, Jianru Xue
Knowl. Based Syst.2
2021 MagDR: Mask-Guided Detection and Reconstruction for Defending Deepfakes
abstract
1Deepfakes raised serious concerns on the authenticity of visual contents. Prior works revealed the possibility to disrupt deepfakes by adding adversarial perturbations to the source data, but we argue that the threat has not been eliminated yet. This paper presents MagDR, a mask-guided detection and reconstruction pipeline for defending deepfakes from adversarial attacks. MagDR starts with a detection module that defines a few criteria to judge the abnormality of the output of deepfakes, and then uses it to guide a learnable reconstruction procedure. Adaptive masks are extracted to capture the change in local facial regions. In experiments, MagDR defends three main tasks of deepfakes, and the learned reconstruction pipeline transfers across input data, showing promising performance in defending both black-box and white-box attacks.
Lingxi Xie, Shanmin Pang, Bo Zhang 0010
CVPR3
2021 Appending Adversarial Frames for Universal Video Attack
abstract
This paper investigates the problem of generating adversarial examples for video classification. We project all videos onto a semantic space and a perception space, and point out that adversarial attack is to find a counterpart which is close to the target in the perception space but far from the target in the semantic space. Based on this formulation, we notice that conventional attacking methods mostly used Euclidean distance to measure the perception space, but we propose to make full use of the property of videos and assume a modified video with a few consecutive frames replaced by dummy contents (e.g., a black frame with texts of `thank you for watching' on it) to be close to the original video in the perception space though they have a large Euclidean gap. This leads to a new attack approach which only adds perturbations on the newly-added frames. We show its high success rates in attacking six state-of-the-art video classification networks, as well as its universality, i.e., transferring well across videos and models.
Lingxi Xie, Shanmin Pang, Qi Tian 0001
WACV3
2021 Deep Subspace Mutual Learning for cancer subtypes prediction
abstract
MOTIVATION: Precise prediction of cancer subtypes is of significant importance in cancer diagnosis and treatment. Disease etiology is complicated existing at different omics levels; hence integrative analysis provides a very effective way to improve our understanding of cancer. RESULTS: We propose a novel computational framework, named Deep Subspace Mutual Learning (DSML). DSML has the capability to simultaneously learn the subspace structures in each available omics data and in overall multi-omics data by adopting deep neural networks, which thereby facilitates the subtype's prediction via clustering on multi-level, single-level and partial-level omics data. Extensive experiments are performed in five different cancers on three levels of omics data from The Cancer Genome Atlas. The experimental analysis demonstrates that DSML delivers comparable or even better results than many state-of-the-art integrative methods. AVAILABILITY AND IMPLEMENTATION: An implementation and documentation of the DSML is publicly available at https://github.com/polytechnicXTT/Deep-Subspace-Mutual-Learning.git.
Bo Yang 0041, Tingting Xin, Shanmin Pang, Meng Wang 0001
Bioinform.3
2021 Multi-view spectral clustering via common structure maximization of local and global representations
Wenyu Hao, Shanmin Pang
Neural Networks2
2021 Integrating Multi-Omic Data With Deep Subspace Fusion Clustering for Cancer Subtype Prediction
abstract
One type of cancer usually consists of several subtypes with distinct clinical implications, thus the cancer subtype prediction is an important task in disease diagnosis and therapy. Utilizing one type of data from molecular layers in biological system to predict is difficult to bridge the cancer genome to cancer phenotypes, since the genome is neither simple nor independent but rather complicated and dysregulated from multiple molecular mechanisms. Similarity Network Fusion (SNF) has been recently proposed to integrate diverse omics data for improving the understanding of tumorigenesis. SNF adopts Euclidean distance to measure the similarity between patients, which shows some limitations. In this article, we introduce a novel prediction technique as an extension of SNF, namely Deep Subspace Fusion Clustering (DSFC). DSFC utilizes auto-encoder and data self-expressiveness approaches to guide a deep subspace model, which can achieve effective expression of discriminative similarity between patients. As a result, the dissimilarity between inter-cluster is delivered and enhanced compactness of intra-cluster is achieved at the same time. The validity of DSFC is examined by extensive simulations over six different cancer through three levels omics data. The survival analysis demonstrates that DSFC delivers comparable or even better results than many state-of-the-art integrative methods.
Bo Yang 0041, Shanmin Pang, Xuequn Shang 0001, Xueqing Zhao, Minghui Han
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Label Enhancement with Sample Correlations via Low-Rank Representation
abstract
Compared with single-label and multi-label annotations, label distribution describes the instance by multiple labels with different intensities and accommodates to more-general conditions. Nevertheless, label distribution learning is unavailable in many real-world applications because most existing datasets merely provide logical labels. To handle this problem, a novel label enhancement method, Label Enhancement with Sample Correlations via low-rank representation, is proposed in this paper. Unlike most existing methods, a low-rank representation method is employed so as to capture the global relationships of samples and predict implicit label correlation to achieve label enhancement. Extensive experiments on 14 datasets demonstrate that the algorithm accomplishes state-of-the-art results as compared to previous label enhancement baselines.
Haoyu Tang 0002, Jihua Zhu, Qinghai Zheng, Jun Wang 0024, Shanmin Pang, Zhongyu Li 0002
AAAI5
2020 Meta Generalized Network for Few-Shot Classification
abstract
Few-shot classification aims to learn a well generalized model with very limited labeled examples. There are mainly two directions for this aim, namely, meta- and metric-learning. Meta learning trains models in a particular way to fast adapt to new tasks, but it neglects variational features of images. Metric learning considers relationships among same or different classes, however on the downside, it usually fails to achieve competitive performance on unseen boundary examples. In this paper, we propose a Meta Generalized Network (MGNet) that aims to combine advantages of both meta- and metric-learning. There are two novel components in MGNet. Specifically, we first develop a meta backbone training method that learns a flexible feature extractor and a classifier initializer efficiently, delightedly leading to fast adaption to unseen few-shot tasks without overfitting. Second, we design a trainable adaptive interval model to improve the cosine classifier, which increases the recognition accuracy of hard examples. We train the meta backbone in the training stage by all classes, and fine-tune the meta-backbone as well as train the adaptive classifier in the testing stage. We evaluate MGNet on three standard image recognition benchmarks, and experimental results validate the superiority over recent competitive methods.
Shanmin Pang, Yaochen Li
ICPR2
2020 Look, Listen and Infer
abstract
Inspired by the ability of human beings on recognizing the relations between visual scenes and sounds, many cross-modal learning methods have been developed for modeling images or videos and associated sounds. In this work, for the first time, a Look, Listen and Infer Network (LLINet) is proposed to learn a zero-shot model that can infer the relations of visual scenes and sounds from novel categories never appeared before. LLINet is mainly desired to qualify for two tasks, i.e., image-audio cross-modal retrieval and sound localization in each image. Towards this end, it is designed as a two-branch encoding network that builds a common space for images and audios. Besides, a cross-modal attention mechanism is proposed in LLINet to localize sound objects. To evaluate LLINet, a new data set, named INSTRUMENT-32CLASS, is collected in this work. Besides zero-shot cross-modal retrieval and sound localization, a zero-shot image recognition task based on sounds is also conducted on this database. All experimental results on these tasks demonstrate the effectiveness of LLINet, indicating that zero-shot learning for visual scenes and sounds is feasible. The project page for LLINet is available at https://llinet.github.io/.
Ruijian Jia, Shanmin Pang, Jihua Zhu, Jianru Xue
ACM Multimedia3
2020 Spatial-Content Image Search in Complex Scenes
abstract
Although the topic of image search has been heavily studied in the last two decades, many works have focused on either instance-level retrieval or semantic-level retrieval. In this work, we develop a novel visually similar spatial-semantic method, namely spatial-content image search, to search images that not only share the same spatial-semantics but also enjoy visual consistency as the query image in complex scenes. We achieve the goal by capturing spatial-semantic concepts as well as the visual representation of each concept contained in an image. Specifically, we first generate a set of bounding boxes and their category labels representing spatial-semantic constraints with YOLOV3, and then obtain visual content of each bounding box with deep features extracted from a convolutional neural network. After that, we customize a similarity computation method that evaluates the relevance between dataset images and input queries according to the developed image representations. Experimental results on two large-scale benchmark retrieval datasets with images consisting of multiple objects demonstrate that our method provides an effective way to query image databases. Our code is available at https://github.com/MaJinWakeUp/spatial-content.
Shanmin Pang, Bo Yang 0041, Jihua Zhu, Yaochen Li
WACV2
2020 Feature concatenation multi-view subspace clustering
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Shanmin Pang, Jun Wang 0024, Yaochen Li
Neurocomputing4
2020 Constrained bilinear factorization multi-view subspace clustering
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Shanmin Pang, Xiuyi Jia
Knowl. Based Syst.5
2020 Simultaneously merging multi-robot grid maps at different resolutions
Zutao Jiang, Jihua Zhu, Congcong Jin, Yiqiong Zhou, Shanmin Pang
Multim. Tools Appl.6
2020 Unsupervised semantic-based convolutional features aggregation for image retrieval
Shanmin Pang, Jihua Zhu, Lin Wang 0026
Multim. Tools Appl.2
2020 Self-Weighting and Hypergraph Regularization for Multi-view Spectral Clustering
abstract
Leveraging the consensus and complementary principle to find a common representation for different views is an essential problem of multi-view clustering. To address the problem, many Low-Rank Representation (LRR) based methods have been proposed. However, existing LRR based methods have two common limitations: 1) they adopt graph regularization that only considers simple pairwise similarities among data points, and 2) they do not generally characterize the importance of each view. In this letter, we correspondingly utilize hypergraph regularization and a self-weighting strategy to handle the limitations with an LRR based model. Specifically, in our model, we construct hypergraph Laplacian matrices of each view that explicitly contain high order relations among data points, to improve the usage of complementary information. Meanwhile, the self-weighting strategy that preserves view specific information and assigns adaptive weights to each view is leveraged to take full advantage of multi-view consensus information. Based on the Augmented Lagrangian Multiplier (ALM) scheme, we design an effective alternating iterative strategy to optimize the model. Extensive experiments conducted on four benchmark datasets validate the superiority of our method.
Wenyu Hao, Shanmin Pang, Jihua Zhu, Yaochen Li
IEEE Signal Process. Lett.2
2020 Registration of Multi-View Point Sets Under the Perspective of Expectation-Maximization
abstract
Multi-view registration plays a critical role in 3D model reconstruction. To solve this problem, most previous methods align point sets by either partially exploring available information or blindly utilizing unnecessary information, which may lead to undesired results or extra computation complexity. Accordingly, we propose a novel solution for the multi-view registration under the perspective of Expectation-Maximization (EM). The proposed method assumes that each data point is generated from one unique Gaussian Mixture Model (GMM), where its corresponding points in other point sets are regarded as Gaussian centroids with equal covariance and membership probabilities. As it is difficult to obtain real corresponding points in the registration problem, they are approximated by the nearest neighbor in each other aligned point sets. Based on this assumption, it is reasonable to define the likelihood function including all rigid transformations, which require to be estimated for multi-view registration. Subsequently, the EM algorithm is derived to estimate rigid transformations with one Gaussian covariance by maximizing the likelihood function. Since the GMM component number is automatically determined by the number of point sets, there is no trade-off between registration accuracy and efficiency in the proposed method. Finally, the proposed method is tested on several benchmark data sets and compared with state-of-the-art algorithms. Experimental results demonstrate its superior performance on the accuracy, efficiency, and robustness for multi-view registration.
Jihua Zhu, Zhongyu Li 0002, Shanmin Pang
IEEE Trans. Image Process.5
2019 Jointly Detecting and Retrieving Vehicles from Road Image Sequences based on CNN
abstract
In this paper, a CNN-based vehicle detection and retrieval framework is proposed for the intelligent transportation system. Firstly, the vehicle target is detected from the traffic scene. The proposed object detection method uses a fully convolutional neural network (CNN) based on SqueezeNet, which has the characteristics of real-time, high accuracy and has small model size. Secondly, an intra-class image retrieval method is presented to search vehicles which are similar to the target vehicle in the dataset. The image retrieval results can be used for traffic scenes simulation and modeling. The experiments and comparisons prove the effectiveness of our framework.
Yaochen Li, Yuehu Liu, Shanmin Pang, Le Wang 0003, Huihui Huo
IV4
2019 Multi-view registration based on weighted LRS matrix decomposition of motions
abstract
Recently, the low‐rank and sparse (LRS) matrix decomposition has been introduced as an effective mean to solve the multi‐view registration. It views each available relative motion as a block element to reconstruct one sparse matrix, which then is used to approximate the low‐rank matrix, where global motions can be recovered for multi‐view registration. However, this approach is sensitive to the sparsity of the reconstructed matrix and it treats all block elements equally in spite of their varied reliabilities. Therefore, this study proposes an effective approach for multi‐view registration by weighted LRS matrix decomposition. On the basis of the inverse symmetry property of relative motions, it first proposes a completion method to reduce the sparsity of the reconstructed matrix. The reduced sparsity of the reconstructed matrix can improve the robustness and efficiency of LRS matrix decomposition. Then, it proposes the weighted LRS matrix decomposition, where each block element is assigned with one estimated weight to denote its reliability. By introducing the weight, more accurate registration results can be efficiently recovered from the estimated low‐rank matrix. Experimental results tested on public datasets illustrate the superiority of the proposed approach over the state‐of‐the‐art approaches on robustness, accuracy and efficiency.
Congcong Jin, Jihua Zhu, Yaochen Li, Shanmin Pang, Lei Chen 0011, Jun Wang 0024
IET Comput. Vis.4
2019 Efficient registration of multi-view point sets by K-means clustering
Jihua Zhu, Zutao Jiang, Georgios Evangelidis 0002, Changqing Zhang 0002, Shanmin Pang, Zhongyu Li 0002
Inf. Sci.5
2019 Co-weighting semantic convolutional features for object retrieval
Jihua Zhu, Shanmin Pang, Weili Guan, Zhongyu Li 0002, Yaochen Li, Xueming Qian
J. Vis. Commun. Image Represent.3
2019 Unifying Sum and Weighted Aggregations for Efficient Yet Effective Image Representation Computation
abstract
Embedding and aggregating a set of local descriptors (e.g. SIFT) into a single vector is normally used to represent images in image search. Standard aggregation operations include sum and weighted aggregations. While showing high efficiency, sum aggregation lacks discriminative power. In contrast, weighted aggregation shows promising retrieval performance but suffers extremely high time cost. In this work, we present a general mixed aggregation method that unifies sum and weighted aggregation methods. Owing to its general formulation, our method is able to balance the trade-off between retrieval quality and image representation efficiency. Additionally, to improve query performance, we propose computing multiple weighting coefficients rather than one for each to be aggregated vector by partitioning them into several components with negligible computational cost. Extensive experimental results on standard public image retrieval benchmarks demonstrate that our aggregation method achieves state-of-the-art performance while showing over ten times speedup over baselines.
Shanmin Pang, Jianru Xue, Jihua Zhu, Li Zhu 0003, Qi Tian 0001
IEEE Trans. Image Process.1
2019 Deep Feature Aggregation and Image Re-Ranking With Heat Diffusion for Image Retrieval
abstract
Image retrieval based on deep convolutional features has demonstrated state-of-the-art performance in popular benchmarks. In this paper, we present a unified solution to address deep convolutional feature aggregation and image re-ranking by simulating the dynamics of heat diffusion. A distinctive problem in image retrieval is that repetitive or bursty features tend to dominate final image representations, resulting in representations less distinguishable. We show that by considering each deep feature as a heat source, our unsupervised aggregation method is able to avoid over-representation of bursty features. We additionally provide a practical solution for the proposed aggregation method and further show the efficiency of our method in experimental evaluation. Inspired by the aforementioned deep feature aggregation method, we also propose a method to re-rank a number of top ranked images for a given query image by considering the query as the heat source. Finally, we extensively evaluate the proposed approach with pre-trained and fine-tuned deep networks on common public benchmarks and show superior performance compared to previous work.
Shanmin Pang, Jianru Xue, Jihua Zhu, Vicente Ordonez
IEEE Trans. Multim.1
2019 Improving Object Retrieval Quality by Integration of Similarity Propagation and Query Expansion
abstract
Re-ranking is an essential step for accurate image retrieval, due to its well-known power in performance improvement. Although numerous works have been proposed for re-ranking, many of them are only customized for a certain image representation model. In contrast to most existing techniques, we develop generalized re-ranking algorithms that are applicable to different kinds of image encodings in this paper. We first employ a quite successful theory of similarity propagation to reconstruct vectors of a query and its top ranked images and, subsequently, get a re-ranked list by comparing the new image vectors. Furthermore, considering that the just mentioned strategy is directly compatible with query expansion and, thus, in order to leverage advantages of this milestone, we then propose integrating them into a unified framework for maximizing re-ranking benefits. Our re-ranking algorithms are memory and computation efficient, and experimental results on benchmark datasets demonstrate that they compare favorably with the state of the art. Our code is available at https://github.com/MaJinWakeUp/rerank.
Shanmin Pang, Jihua Zhu, Jianru Xue, Qi Tian 0001
IEEE Trans. Multim.1
2018 Deep Subspace Similarity Fusion for the Prediction of Cancer Subtypes
Bo Yang 0041, Shuhui Liu, Shanmin Pang, Chenpai Pang, Xuequn Shang 0001
BIBM3
2018 Adaptive Co-Weighting Deep Convolutional Features for Object Retrieval
abstract
Aggregating deep convolutional features into a global image vector has attracted sustained attention in image retrieval. In this paper, we propose an efficient unsupervised aggregation method that uses an adaptive Gaussian filter and an element-value sensitive vector to co-weight deep features. Specifically, the Gaussian filter assigns large weights to features of region-of-interests (RoI) by adaptively determining the RoI's center, while the element-value sensitive channel vector suppresses burstiness phenomenon by assigning small weights to feature maps with large sum values of all locations. Experimental results on benchmark datasets validate the proposed two weighting schemes both effectively improve the discrimination power of image vectors. Furthermore, with the same experimental setting, our method outperforms other very recent aggregation approaches by a considerable margin.
Jihua Zhu, Shanmin Pang, Zhongyu Li 0002, Yaochen Li, Xueming Qian
ICME3
2018 Building discriminative CNN image representations for object retrieval using the replicator equation
Shanmin Pang, Jihua Zhu, Vicente Ordonez, Jianru Xue
Pattern Recognit.1
2018 Large-scale vocabularies with local graph diffusion and mode seeking
Shanmin Pang, Jianru Xue, Zhanning Gao, Lihong Zheng, Li Zhu 0003
Signal Process. Image Commun.1
2017 Isometric hashing for image retrieval
Bo Yang 0041, Xuequn Shang 0001, Shanmin Pang
Signal Process. Image Commun.3
2016 Multiple Cartesian K-Medoids for a Fine Quantization
abstract
K-means is a widely used method for the process of vector quantization in image retrieval, and its results will directly affect the subsequent retrieval quality. Although k-means is popular in image retrieval, it has some obvious disadvantages, such as randomness and sensitivity to outliers. This paper presents a new model, namely Multiple Cartesian K-medoids, to replace k-means for quantization and retrieval. The proposed model proceeds in two steps. The first step is to establish multiple K-medoids model to finely quantize feature vectors to codewords. Then, the second step establishes local linear search: adopt an inverted file for efficiently searching candidate nearest neighbors of a given query, and finally obtains accurate neighbors of the query by re-ranking these candidate neighbors with Euclidean distances of the original feature vectors. Experimental results show that the proposed method is effective, and substantially improves the search accuracy of the returned nearest neighbors.
Lihua Tian, Shanmin Pang, Chen Li 0033
ICPADS2
2016 Democratic Diffusion Aggregation for Image Retrieval
abstract
Content-based image retrieval is an important research topic in the multimedia field. In large-scale image search using local features, image features are encoded and aggregated into a compact vector to avoid indexing each feature individually. In the aggregation step, sum-aggregation is wildly used in many existing works and demonstrates promising performance. However, it is based on a strong and implicit assumption that the local descriptors of an image are identically and independently distributed in descriptor space and image plane. To address this problem, we propose a new aggregation method named democratic diffusion aggregation (DDA) with weak spatial context embedded. The main idea of our aggregation method is to re-weight the embedded vectors before sum-aggregation by considering the relevance among local descriptors. Different from previous work, by conducting a diffusion process on the improved kernel matrix, we calculate the weighting coefficients more efficiently without any iterative optimization. Besides considering the relevance of local descriptors from different images, we also discuss an efficient query fusion strategy which uses the initial top-ranked image vectors to enhance the retrieval performance. Experimental results show that our aggregation method exhibits much higher efficiency (about × 14 faster) and better retrieval accuracy compared with previous methods, and the query fusion strategy consistently improves the retrieval quality.
Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001
IEEE Trans. Multim.4
2015 Fast Democratic Aggregation and Query Fusion for Image Search
abstract
In image search using local features, to avoid indexing each feature individually, encoding methods are popularly adopted to embed and aggregate local features of an image into a compact vector. Democratic aggregation with triangulation embedding (T-embedding) exhibits significant retrieval accuracy improvement over previous works. However, it suffers high computational complexity. To address this problem and consistently improve the retrieval performance, we propose a new democratic method to accelerate aggregating step without accuracy lost. We also embed weak spatial context in the kernel construction to depress co-occurrence caused by local feature detector. Furthermore, we enhance the retrieval performance with an efficient query fusion strategy. The evaluation on public datasets shows that our democratic aggregation is an order of magnitude faster than the original democratic aggregation with comparable retrieval accuracy, and the query fusion achieves a significant accuracy improvement over previous works.
Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001
ICMR4
2015 Image re-ranking with an alternating optimization
Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001
Neurocomputing1
2014 Image Re-ranking with an Alternating Optimization
abstract
In this work, we propose an efficient image re-ranking method, without additional memory cost compared with the baseline method~\cite{philbin2007object}, to re-rank all retrieved images. The motivation of the proposed method is that, there are usually many visual words in the query image that only give votes to irrelevant images. With this observation, we propose to only use visual words which can help to find relevant images to re-rank the retrieved images. To achieve the goal, we first find some similar images to the query by maximizing a quadratic function when given an initial ranking of the retrieved images. Then we select query visual words with an alternating optimization strategy: (1) at each iteration, select words based on the similar images that we have found and (2) in turn, update the similar images with the selected words. These two steps are repeated until convergence. Experimental results on standard benchmark datasets show that the proposed method outperforms spatial based re-ranking methods.
Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001
ACM Multimedia1
2014 Exploiting local linear geometric structure for identifying correct matches
Shanmin Pang, Jianru Xue, Qi Tian 0001, Nanning Zheng 0001
Comput. Vis. Image Underst.1
2013 Locality preserving verification for image search
abstract
Establishing correct correspondences between two images has a wide range of applications, such as 2D and 3D registration, structure from motion, and image retrieval. In this paper, we propose a new matching method based on spatial constraints. The proposed method has linear time complexity, and is efficient when applying it to image retrieval. The main assumption behind our method is that, the local geometric structure among a feature point and its neighbors, is not easily affected by both geometric and photometric transformations, and thus should be preserved in their corresponding images. We model this local geometric structure by linear coefficients that reconstruct the point from its neighbors. The method is flexible, as it can not only estimate the number of correct matches between two images efficiently, but also determine the correctness of each match accurately. Furthermore, it is simple and easy to be implemented. When applying the proposed method on re-ranking images in an image search engine, it outperforms the-state-of-the-art techniques.
Shanmin Pang, Jianru Xue, Nanning Zheng 0001, Qi Tian 0001
ACM Multimedia1
2012 Large-Scale Bundle Adjustment by Parameter Vector Partition
Shanmin Pang, Jianru Xue, Le Wang 0003, Nanning Zheng 0001
ACCV (4)1