Weilong Peng

dblp:175/1344 · DBLP profile ↗
← Back
43ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0001-5820-889XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 16 since 2021Computer networks · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks
abstract
Most adversarial attacks on point clouds perturb a large number of points, causing widespread geometric changes and limiting applicability in real-world scenarios. While recent works explore sparse attacks by modifying only a few points, such approaches often struggle to maintain effectiveness due to the limited influence of individual perturbations. In this paper, we propose SCP, a sparse and cooperative perturbation framework that selects and leverages a compact subset of points whose joint perturbations produce amplified adversarial effects. Specifically, SCP identifies the subset where the misclassification loss is locally convex with respect to their joint perturbations, determined by checking the positive-definiteness of the corresponding Hessian block. The selected subset is then optimized to generate high-impact adversarial examples with minimal modifications. Extensive experiments show that SCP achieves 100% attack success rates, surpassing state-of-the-art sparse attacks, and delivers superior imperceptibility to dense attacks with far fewer modifications.
Keke Tang, Tianyu Hao, Weilong Peng, Denghui Zhang 0001, Peican Zhu, Zhihong Tian 0001
AAAI4
2026 End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language Models
abstract
Knowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines.
Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang
AAAI4
2026 WaveSculpt: Text-to-3D generation with wavelet-guided score distillation
Weilong Peng, Jianhui Huo, Keke Tang, Yangtao Wang, Yan Wang 0022, Meie Fang
Comput. Aided Geom. Des.1
2026 Transferable and undefendable point cloud attacks via medial axis transform
Keke Tang, Yuze Gao, Weilong Peng, Meie Fang, Peican Zhu
Comput. Aided Geom. Des.3
2026 Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002
Pattern Recognit.5
2026 PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002
Pattern Recognit.9
2026 High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream Transformers
abstract
Recently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git .
Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Imperceptible 3D Point Cloud Attacks on Lattice-based Barycentric Coordinates
abstract
Imperceptible adversarial attacks on 3D point clouds rely on effective constraints. While manifold constraints have notable advantages over Euclidean ones, the global parameterization used in current methods often fails to fully preserve manifold properties. In this paper, we propose to constrain lattice-based barycentric coordinates during attacks from a local parametric perspective to ensure imperceptibility. Specifically, we utilize a permutohedral lattice to partition point clouds into multiple cells, and then extract barycentric coordinates for each point within these cells, forming a local parametric representation of the point clouds. By enforcing local parametric constraints that minimize the displacement of barycentric coordinates, we largely preserve the manifold properties, ultimately leading to improved imperceptibility. Extensive experiments validate that integrating these local parametric constraints into conventional adversarial attacks yields superior imperceptibility, outperforming state-of-the-art methods.
Keke Tang, Ziyong Du, Weilong Peng, Daizong Liu, Ligang Liu 0001, Zhihong Tian 0001
AAAI3
2025 Simplification Is All You Need against Out-of-Distribution Overconfidence
abstract
Deep neural networks (DNNs) often exhibit out-of-distribution (OOD) overconfidence, producing overly confident predictions on OOD samples. We attribute this issue to the inherent over-complexity of DNNs and investigate two key aspects: capacity and nonlinearity. First, we demonstrate that reducing model capacity through knowledge distillation can effectively mitigate OOD overconfidence. Second, we show that selectively reducing nonlinearity by removing ReLU operations further alleviates the issue. Building on these findings, we present a practical guide to model simplification, combining both strategies to significantly reduce OOD overconfidence. Extensive experiments validate the effectiveness of this approach in mitigating OOD overconfidence and demonstrate its superiority over state-of-the-art methods. Additionally, our simplification strategies can be combined with existing OOD detection techniques to further enhance OOD detection performance.
Keke Tang, Weilong Peng, Zhize Wu, Yongwei Nie, Wenping Wang 0001, Zhihong Tian 0001
CVPR3
2025 Imperceptible Adversarial Attacks on Point Clouds Guided by Point-to-Surface Field
abstract
Adversarial attacks on point clouds are crucial for assessing and improving the adversarial robustness of 3D deep learning models. Traditional solutions strictly limit point displacement during attacks, making it challenging to balance imperceptibility with adversarial effectiveness. In this paper, we attribute the inadequate imperceptibility of adversarial attacks on point clouds to deviations from the underlying surface. To address this, we introduce a novel point-to-surface (P2S) field that adjusts adversarial perturbation directions by dragging points back to their original underlying surface. Specifically, we use a denoising network to learn the gradient field of the logarithmic density function encoding the shape’s surface, and apply a distance-aware adjustment to perturbation directions during attacks, thereby enhancing imperceptibility. Extensive experiments show that adversarial attacks guided by our P2S field are more imperceptible, outperforming state-of-the-art methods.
Keke Tang, Weiyao Ke, Weilong Peng, Ziyong Du, Zhize Wu, Peican Zhu, Zhihong Tian 0001
ICASSP3
2025 From Pixels to Shapes: Generative AI for 2D Images and 3D Models
Jianhui Huo, Shijian Xu, Weilong Peng, Yangtao Wang, Yan Wang 0022, Meie Fang
ICIC (19)4
2025 HGACF: A Homogeneous Neighbor Graph Contrastive Learning Framework for Enhanced Collaborative Filtering
Yan Wang 0022, Jinting Nie, Weilong Peng
ICIC (19)4
2025 Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text Retrieval
abstract
Image-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA.
Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016
ICME7
2025 EOOD: Entropy-based Out-of-distribution Detection
abstract
Deep neural networks (DNNs) often exhibit overconfidence when encountering out-of-distribution (OOD) samples, posing significant challenges for deployment. Since DNNs are trained on in-distribution (ID) datasets, the information flow of ID samples through DNNs inevitably differs from that of OOD samples. In this paper, we propose an Entropy-based Out-Of-distribution Detection (EOOD) framework. EOOD first identifies specific block where the information flow differences between ID and OOD samples are more pronounced, using both ID and pseudo-OOD samples. It then calculates the conditional entropy on the selected block as the OOD confidence score. Comprehensive experiments conducted across various ID and OOD settings demonstrate the effectiveness of EOOD in OOD detection and its superiority over state-of-the-art methods.
Guide Yang, Weilong Peng, Yongwei Nie, Peican Zhu, Keke Tang
IJCNN3
2025 EIA: Edge-Aware Imperceptible Adversarial Attacks on 3D Point Clouds
Zhensu Wang, Weilong Peng, Le Wang 0008, Zhizhe Wu, Peican Zhu, Keke Tang
MMM (1)2
2025 MeshPAD: Payload-aware mesh distortion for 3D steganography based on geometric deep learning
Weilong Peng, Keke Tang, Weixuan Tang 0002, Yong Su 0003, Meie Fang, Ping Li 0016
Expert Syst. Appl.1
2025 Adaptive Multi-Lens Phase Modulation for Scale-Aware Privacy-Preserving Human Pose Recognition
abstract
Recently, optical privacy protection has emerged as a promising approach for safeguarding visual privacy at the physical acquisition stage. However, existing methods often face a trade‐off between privacy strength and human pose recognition accuracy, particularly in long‐range and multi‐scale scenarios. To address this challenge, we propose a novel adaptive optical privacy‐preserving framework that integrates a learnable optical modulation system with a human pose recognition network. The core of our method lies in a sparse‐weighted multi‐lens model, where a lightweight multilayer perceptron (MLP) predicts a sparse set of coefficients to linearly combine predefined lens phase profiles based on facial region geometry. This enables dynamic control over the point spread function (PSF), adapting the degree of image degradation to subject scale in real time. Additionally, we introduce a privacy‐aware loss function that selectively reduces facial localization accuracy while preserving body pose information. Extensive experiments on MSCOCO and FLIC datasets demonstrate that the proposed method achieves a favorable balance between privacy protection and pose estimation, outperforming previous optical‐ and software‐based baselines.
Weilong Peng, Quanwei Deng, Mingjie Li 0004, Yangtao Wang, Yan Wang 0022, Lisheng Fan, Meie Fang
IET Softw.1
2025 Dual-Detector Reoptimization for Federated Weakly Supervised Video Anomaly Detection via Adaptive Dynamic Recursive Mapping
abstract
Federated weakly supervised video anomaly detection represents a significant advancement in privacy-preserving collaborative learning, enabling distributed clients to train anomaly detectors using only video-level annotations. However, the inherent challenges of optimizing noisy representation with coarse-grained labels often result in substantial local model errors, which are exacerbated during federated aggregation, particularly in heterogeneous scenarios. To address these limitations, we propose a novel dual-detector framework incorporating adaptive dynamic recursive mapping, which significantly enhances local model accuracy and robustness against representation noise. Our framework integrates two complementary components: a channel-averaged anomaly detector and a channel-statistical anomaly detector, which interact through cross-detector adaptive decision parameters to enable iterative optimization and stable anomaly scoring across all instances. Furthermore, we introduce the scene-similarity adaptive local aggregation algorithm, which dynamically aggregates and learns private models based on scene similarity, thereby enhancing generalization capabilities across diverse scenarios. Extensive experiments conducted on the NVIDIA Jetson AGX Xavier platform using the ShanghaiTech and UBnormal datasets demonstrate the superior performance of our approach in both centralized and federated settings. Notably, in federated environments, our method achieves remarkable improvements of 6.2% and 12.3% in AUC compared to state-of-the-art methods, underscoring its effectiveness in resource-constrained scenarios and its potential for real-world applications in distributed video surveillance systems.
Yong Su 0003, Jiahang Li 0004, Simin An, Hengpeng Xu, Weilong Peng
IEEE Trans. Ind. Informatics5
2025 Continuous Bijection Supervised Pyramid Diffeomorphic Deformation for Learning Tooth Meshes From CBCT Images
abstract
Accurate and high-quality tooth mesh generation from cone-beam computerized tomography (CBCT) is an essential computer-aided technology for digital dentistry. However, existing segmentation-based methods require complicated post-processing and significant manual correction to generate regular tooth meshes. In this paper, we propose a method of continuous bijection supervised pyramid diffeomorphic deformation (PDD) for learning tooth meshes, which could be used to directly generate high-quality tooth meshes from CBCT Images. Overall, we adopt a classic two-stage framework. In the first stage, we devise an enhanced detector to accurately locate and crop every tooth. In the second stage, a PDD network is designed to deform a sphere mesh from low resolution to high one according to pyramid flows based on diffeomorphic mesh deformations, so that the generated mesh approximates the ground truth infinitely and efficiently. To achieve that, a novel continuous bijection distance loss on the diffeomorphic sphere is also designed to supervise the deformation learning, which overcomes the shortcoming of loss based on nearest-neighbour mapping and improves the fitting precision. Experiments show that our method outperforms the state-of-the-art methods in terms of both different evaluation metrics and the geometry quality of reconstructed tooth surfaces.
Zechu Zhang, Weilong Peng, Jinyu Wen, Keke Tang, Meie Fang, David Dagan Feng, Ping Li 0016
IEEE Trans. Multim.2
2024 Manifold Constraints for Imperceptible Adversarial Attacks on Point Clouds
abstract
Adversarial attacks on 3D point clouds often exhibit unsatisfactory imperceptibility, which primarily stems from the disregard for manifold-aware distortion, i.e., distortion of the underlying 2-manifold surfaces. In this paper, we develop novel manifold constraints to reduce such distortion, aiming to enhance the imperceptibility of adversarial attacks on 3D point clouds. Specifically, we construct a bijective manifold mapping between point clouds and a simple parameter shape using an invertible auto-encoder. Consequently, manifold-aware distortion during attacks can be captured within the parameter space. By enforcing manifold constraints that preserve local properties of the parameter shape, manifold-aware distortion is effectively mitigated, ultimately leading to enhanced imperceptibility. Extensive experiments demonstrate that integrating manifold constraints into conventional adversarial attack solutions yields superior imperceptibility, outperforming the state-of-the-art methods.
Keke Tang, Weilong Peng, Jianpeng Wu, Yawen Shi, Daizong Liu, Pan Zhou 0001, Wenping Wang 0001, Zhihong Tian 0001
AAAI3
2024 Image-text Retrieval with Main Semantics Consistency
abstract
Image-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC.
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang
CIKM7
2024 CORES: Convolutional Response-based Score for Out-of-distribution Detection
abstract
Deep neural networks (DNNs) often display overconfidence when encountering out-of-distribution (OOD) samples, posing significant challenges in real-world applications. Capitalizing on the observation that responses on convolutional kernels are generally more pronounced for in-distribution (ID) samples than for OOD ones, this paper proposes the COnvolutional REsponse-based Score (CORES) to exploit these discrepancies for OOD detection. Initially, CORES delves into the extremities of convolutional responses by considering both their magnitude and the frequency of significant values. Moreover, through backtracking from the most prominent predictions, CORES effectively pinpoints sample-relevant kernels across different layers. These kernels, which exhibit a strong correlation to input samples, are integral to CORES's OOD detection capability. Comprehensive experiments across various ID and OOD settings demonstrate CORES's effectiveness in OOD detection and its superiority to the state-of-the-art methods.
Keke Tang, Weilong Peng, Runnan Chen, Peican Zhu, Wenping Wang 0001, Zhihong Tian 0001
CVPR3
2024 FLAT: Flux-Aware Imperceptible Adversarial Attacks on 3D Point Clouds
Keke Tang, Lujie Huang, Weilong Peng, Daizong Liu, Ligang Liu 0001, Zhihong Tian 0001
ECCV (6)3
2024 Reparameterization Head for Efficient Multi-Input Networks
abstract
Reparameterization techniques have demonstrated their efficacy in improving the efficiency of deep neural networks. However, their application has been largely confined to single-input network structures, leaving multi-input ones, commonly encountered in real-world applications, largely unexplored. In this paper, we formulate reparameterization head (RepHead), the first framework designed to introduce reparameterization into multi-input neural networks. RepHead compresses multiple inputs into a single input and employs reconstruction operations to recover them, thereby transforming multi-input networks into single-input, multibranch architectures, thereby enabling the application of reparameterization. We demonstrate the usage of RepHead in both image and point cloud domains. Extensive experimental results validate that the integration of RepHead substantially reduces computational overhead and memory requirements while maintaining minimal performance loss.
Keke Tang, Weilong Peng, Peican Zhu, Zhihong Tian 0001
ICASSP3
2024 IE-aware Consistency Losses for Detailed 3D Face Reconstruction from Multiple Images in the Wild
abstract
3D face reconstruction from multiple in-the-wild images in an unsupervised manner poses a significant challenge, primarily due to the pervasive presence of Intrinsic and Extrinsic inconsistencies in facial features. To tackle this, we introduce a novel set of IE-aware consistency losses designed to effectively mitigate these inconsistencies. Our Local Alignment Loss employs neighborhood search techniques to identify and optimize consistent pixel information, thereby reducing intrinsic inconsistencies. In parallel, our Region Subset Selection Loss filters out regions where significant discrepancies exist between the input and reconstructed images, effectively alleviating extrinsic inconsistencies. Extensive experimental results validate the effectiveness of our IE-aware consistency losses in reconstructing detailed 3D facial geometry from images captured in uncontrolled environments.
Weilong Peng, Keke Tang, Kongyang Chen, Yangtao Wang, Ping Li 0016, Meie Fang
ICME1
2024 SymAttack: Symmetry-aware Imperceptible Adversarial Attacks on 3D Point Clouds
abstract
Adversarial attacks on point clouds are crucial for assessing and improving the adversarial robustness of 3D deep learning models. Despite leveraging various geometric constraints, current adversarial attack strategies often suffer from inadequate imperceptibility. Given that adversarial perturbations tend to disrupt the inherent symmetry in objects, we recognize this disruption as the primary cause of the lack of imperceptibility in these attacks. In this paper, we introduce a novel framework, symmetry-aware imperceptible adversarial attacks on 3D point clouds (SymAttack), to address this issue. Our approach starts by identifying part- and patch-level symmetry elements, and grouping points based on semantic and Euclidean distances, respectively. During the adversarial attack iterations, we intentionally adjust the perturbation vectors on symmetric points relative to their symmetry plane. By preserving symmetry within the attack process, SymAttack significantly enhances imperceptibility. Extensive experiments validate the effectiveness of SymAttack in generating imperceptible adversarial point clouds, demonstrating its superiority over the state-of-the-art methods.
Keke Tang, Zhensu Wang, Weilong Peng, Lujie Huang, Le Wang 0008, Peican Zhu, Wenping Wang 0001, Zhihong Tian 0001
ACM Multimedia3
2024 MIT: Multi-cue Injected Transformer for Two-Stage HOI Detection
Weilong Peng, Qingfeng Chen, Keke Tang, Meng Xing, Meie Fang
PRCV (7)1
2024 High-Quality Fusion and Visualization for MR-PET Brain Tumor Images via Multi-Dimensional Features
abstract
The fusion of magnetic resonance imaging and positron emission tomography can combine biological anatomical information and physiological metabolic information, which is of great significance for the clinical diagnosis and localization of lesions. In this paper, we propose a novel adaptive linear fusion method for multi-dimensional features of brain magnetic resonance and positron emission tomography images based on a convolutional neural network, termed as MdAFuse. First, in the feature extraction stage, three-dimensional feature extraction modules are constructed to extract coarse, fine, and multi-scale information features from the source image. Second, at the fusion stage, the affine mapping function of multi-dimensional features is established to maintain a constant geometric relationship between the features, which can effectively utilize structural information from a feature map to achieve a better reconstruction effect. Furthermore, our MdAFuse comprises a key feature visualization enhancement algorithm designed to observe the dynamic growth of brain lesions, which can facilitate the early diagnosis and treatment of brain tumors. Extensive experimental results demonstrate that our method is superior to existing fusion methods in terms of visual perception and nine kinds of objective image fusion metrics. Specifically, in the results of MR-PET fusion, the SSIM (Structural Similarity) and VIF (Visual Information Fidelity) metrics show improvements of 5.61% and 13.76%, respectively, compared to the current state-of-the-art algorithm. Our project is publicly available at: https://github.com/22385wjy/MdAFuse.
Jinyu Wen, Amei Chen, Weilong Peng, Meie Fang, C. L. Philip Chen, Ping Li 0016
IEEE Trans. Image Process.4
2023 Deep Manifold Attack on Point Clouds via Parameter Plane Stretching
abstract
Adversarial attack on point clouds plays a vital role in evaluating and improving the adversarial robustness of 3D deep learning models. Current attack methods are mainly applied by point perturbation in a non-manifold manner. In this paper, we formulate a novel manifold attack, which deforms the underlying 2-manifold surfaces via parameter plane stretching to generate adversarial point clouds. First, we represent the mapping between the parameter plane and underlying surface using generative-based networks. Second, the stretching is learned in the 2D parameter domain such that the generated 3D point cloud fools a pretrained classifier with minimal geometric distortion. Extensive experiments show that adversarial point clouds generated by manifold attack are smooth, undefendable and transferable, and outperform those samples generated by the state-of-the-art non-manifold ones.
Keke Tang, Jianpeng Wu, Weilong Peng, Yawen Shi, Peng Song 0001, Zhaoquan Gu, Zhihong Tian 0001, Wenping Wang 0001
AAAI3
2023 Matching Words for Out-of-distribution Detection
abstract
Deep neural networks often exhibit the overconfidence issue when encountering out-of-distribution (OOD) samples. To address this, leveraging large-scale pre-trained models like CLIP has shown promise. While CLIP has the capability to encode a vast array of interconnected concepts, current OOD detection methods based on it primarily focus on ID categories and a limited set of OOD categories. In this paper, we propose a novel approach that harnesses the power of WordNet to fully exploit the rich knowledge encapsulated within CLIP, resulting in enhanced OOD detection performance. Our methodology involves constructing a word tree that includes both in-distribution (ID) words and a large set of semantically similar OOD words selected from WordNet. By matching a test image with the concepts of the words in the word tree using CLIP, we estimate the probability of the image being classified as either ID or OOD. Furthermore, we introduce a conditional random field model to effectively handle both the parent-child and the sibling-sibling conflicts in the concept matching results. Extensive experiments under various ID/OOD settings demonstrate the effectiveness of our approach and its superiority over state-of-the-art methods.
Keke Tang, Xujian Cai, Weilong Peng, Daizong Liu, Peican Zhu, Pan Zhou 0001, Zhihong Tian 0001, Wenping Wang 0001
ICDM3
2023 OOD Attack: Generating Overconfident out-of-Distribution Examples to Fool Deep Neural Classifiers
abstract
Deep neural networks (DNNs) are dominating various computer vision solutions. However, DNN classifiers suffer from the out-of-distribution (OOD) overconfidence issue, i.e., making overconfident predictions on OOD samples. In this paper, we consider a new OOD attack task, i.e., generating OOD examples that fool DNN classifiers to trap into this issue. Specifically, we first generate seed examples by sampling from common OOD distributions, and then lift the prediction to be overconfident. Extensive experiments with different seeds and confidence-lifting solutions under white-and black-box settings validate the feasibility of OOD attack. Besides, we demonstrate its usefulness in evaluating OOD detection and alleviating the OOD overconfidence issue.
Keke Tang, Xujian Cai, Weilong Peng, Shudong Li, Wenping Wang 0001
ICIP3
2023 Are Deep Point Cloud Classifiers Suffer From Out-of-distribution Overconfidence Issue?
abstract
3D point cloud perception using deep neural networks (DNNs) has been a trend for various application scenarios. However, the black-box nature of DNNs will bring many hidden risks as in the 2D image field. In this paper, we present a preliminary evaluation on the out-of-distribution (OOD) overconfidence issue of deep point cloud classifiers, which has been proven to exist in deep 2D image classifiers, i.e., OOD inputs will lead to overconfident predictions on predefined categories. We also investigate whether a simple thresholding baseline and two modern OOD detection solutions can handle the issue by detecting OOD samples. Extensive experiments with four representative deep point cloud classifiers train/evaluate on different in/out-of-distribution point clouds validate the severity and knottiness of the OOD overconfidence issue. Our investigation will provide the groundwork for future studies on handling the OOD overconfidence issue of DNN classifiers for 3D point clouds.
Keke Tang, Yawen Shi, Weilong Peng, Peican Zhu
SMC5
2023 Rethinking Perturbation Directions for Imperceptible Adversarial Attacks on Point Clouds
abstract
Adversarial attacks have been successfully extended to the field of point clouds. Besides applying the common perturbation guided by the gradient, adversarial attacks on point clouds can be conducted by applying directional perturbations, e.g., along normal and along the tangent plane. In this article, we first investigate whether adversarial attacks with these two orthogonal directional perturbations are more imperceptible than that with the gradient-aware perturbation. Second, we investigate the deeper difference between adversarial attacks with these two directional perturbations, and whether they are applicable to the same scenarios. Third, based on the verification results that the above two directional perturbations have different sensitiveness to curvature, we devise a novel normal-tangent attack (NTA) framework with a hybrid directional perturbation scheme that adaptively chooses the direction according to the curvature of the local shape around the point. Extensive experiments on two publicly available data sets, e.g., ModelNet40 and ShapeNet Part, with classifiers in three representative networks, e.g., PointNet++, DGCNN, PointConv, validate the effectiveness of NTA, and the superiority to the state-of-the-art methods.
Keke Tang, Yawen Shi, Tianrui Lou, Weilong Peng, Peican Zhu, Zhaoquan Gu, Zhihong Tian 0001
IEEE Internet Things J.4
2023 RepPVConv: attentively fusing reparameterized voxel features for efficient 3D point cloud perception
Keke Tang, Weilong Peng, Yanling Zhang, Meie Fang, Zheng Wang 0002, Peng Song 0001
Vis. Comput.3
2021 CODEs: Chamfer Out-of-Distribution Examples against Overconfidence Issue
abstract
Overconfident predictions on out-of-distribution (OOD) samples is a thorny issue for deep neural networks. The key to resolve the OOD overconfidence issue inherently is to build a subset of OOD samples and then suppress predictions on them. This paper proposes the Chamfer OOD examples (CODEs), whose distribution is close to that of in-distribution samples, and thus could be utilized to alleviate the OOD overconfidence issue effectively by suppressing predictions on them. To obtain CODEs, we first generate seed OOD examples via slicing&splicing operations on in-distribution samples from different categories, and then feed them to the Chamfer generative adversarial network for distribution transformation, without accessing to any extra data. Training with suppressing predictions on CODEs is validated to alleviate the OOD overconfidence issue largely without hurting classification accuracy, and outperform the state-of-the-art methods. Besides, we demonstrate CODEs are useful for improving OOD detection and classification.
Keke Tang, Dingruibo Miao, Weilong Peng, Jianpeng Wu, Yawen Shi, Zhaoquan Gu, Zhihong Tian 0001, Wenping Wang 0001
ICCV3
2021 Spatio-temporal multi-factor model for individual identification from biological motion
Yong Su 0003, Weilong Peng, Meng Xing, Zhiyong Feng 0002
Ad Hoc Networks2
2021 VDARN: Video Disentangling Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Yong Su 0003, Meng Xing, Simin An, Weilong Peng, Zhiyong Feng 0002
Ad Hoc Networks4
2021 Disentangling style on dynamic aligned poses for individual identification
Yong Su 0003, Meng Xing, Weilong Peng, Zhiyong Feng 0002
Ad Hoc Networks4
2021 Ventral & Dorsal Stream Theory based Zero-Shot Action Recognition
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng
Pattern Recognit.4
2019 Dynamic hand gesture recognition using motion pattern and shape descriptors
Meng Xing, Jing Hu 0007, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng, Jinqing Zheng
Multim. Tools Appl.5
2018 Sequential Articulated Motion Reconstruction from a Monocular Image Sequence
abstract
In this article, we present a sequential approach for articulated motion estimation from a 2D skeleton sequence. This is a challenging task due to the complexity of human movements and the inherent depth ambiguities. The proposed approach models the human movement on a kinematic manifold with the tangent bundle, which is a natural geometrical representation of articulated motion. Combined with a second-order stochastic dynamic model based on the Markov hypothesis, we generalize the Extended Rauch Tung Striebel smoother to a Riemannian manifold to simulate the process of human movement. The human motor system might violate the Markov hypothesis when the human body is subject to external forces, and therefore a refinement stage is introduced to correct the estimation error. Specifically, the current estimation is refined in a feasible solution region consisting of a set of local estimations. This region is called a simplex, in which each element can be represented by a convex hull of all ingredients. We have proved that the refinement problem can be converted into a convex optimization problem with the simplicial constraint. Since the proposed formulation conforms to the principles of kinematic and spatio-temporal continuity of articulated motion, the reconstruction ambiguity can be alleviated essentially. The performance of the proposed algorithm is conducted on multiple synthetic sequences from the CMU and the HDM05 MoCap databases. The results show that, without requiring any training data, the proposed approach achieves greater accuracy over state-of-the-art baselines. Furthermore, the proposed approach outperforms two baselines on real sequences from the Human3.6m MoCap database.
Yong Su 0003, Zhiyong Feng 0002, Weilong Peng, Meng Xing
ACM Trans. Multim. Comput. Commun. Appl.4
2017 Parametric T-Spline Face Morphable Model for Detailed Fitting in Shape Subspace
abstract
Pre-learnt subspace methods, e.g., 3DMMs, are significant exploration for the synthesis of 3D faces by assuming that faces are in a linear class. However, the human face is in a nonlinear manifold, and a new test are always not in the pre-learnt subspace accurately because of the disparity brought by ethnicity, age, gender, etc. In the paper, we propose a parametric T-spline morphable model (T-splineMM) for 3D face representation, which has great advantages of fitting data from an unknown source accurately. In the model, we describe a face by C^2 T-spline surface, and divide the face surface into several shape units (SUs), according to facial action coding system (FACS), on T-mesh instead of on the surface directly. A fitting algorithm is proposed to optimize coefficients of T-spline control point components along pre-learnt identity and expression subspaces, as well as to optimize the details in refinement progress. As any pre-learnt subspace is not complete to handle the variety and details of faces and expressions, it covers a limited span of morphing. SUs division and detail refinement make the model fitting the facial muscle deformation in a larger span of morphing subspace. We conduct experiments on face scan data, kinect data as well as the space-time data to test the performance of detail fitting, robustness to missing data and noise, and to demonstrate the effectiveness of our model. Convincing results are illustrated to demonstrate the effectiveness of our model compared with the popular methods.
Weilong Peng, Zhiyong Feng 0002, Chao Xu 0003, Yong Su 0003
CVPR1
2016 3D face modeling based on structure optimization and surface reconstruction with B-Spline
Weilong Peng, Chao Xu 0003, Zhiyong Feng 0002
Neurocomputing1