VLDB 2026 Research / reviewers in the wild / expert
Keke Tang
dblp:162/3984
· DBLP profile ↗
71ranked-venue papers
20as first author
66since 2021 · last 2027
0000-0003-0377-1022ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 12 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 16 first-author · 37 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Rethinking 3D point cloud adversarial attacks from models' inherent focus
Xiaowen Cai 0001, Shuqin Chen, Junhao Dong 0001, Keke Tang, Zhongliang Guo 0001, Daizong Liu |
Expert Syst. Appl. | 5 |
| 2026 | Towards Unified Vision-Language Models with Incomplete Multi-Modal InputsabstractVideo-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due to deactivated sensors (e.g., cameras are unavailable due to data privacy), yielding modality-incomplete data and leading to inconsistency between training and testing data. While straightforward incomplete input can boast training generalization-ability and lead to training failure, its potential risks to VLMs regarding safety and trustworthiness have been largely neglected. To this end, we make the first attempt to propose a unified incomplete video-language model to process the incomplete multi-modal inputs. Extensive experimental results show that our method can serve as a plug-and-play module for previous works to improve their performance in various multi-modal tasks. Wanlong Fang, Changshuo Wang 0001, Keke Tang, Daizong Liu, Wei Ji 0008 |
AAAI | 4 |
| 2026 | Transferable Hypergraph Attack via Injecting Nodes into Pivotal HyperedgesabstractRecent studies have demonstrated that hypergraph neural networks (HGNNs) are susceptible to adversarial attacks. However, existing methods rely on the specific information mechanisms of target HGNNs, overlooking the common vulnerability caused by the significant differences in hyperedge pivotality along aggregation paths in most HGNNs, thereby limiting the transferability and effectiveness of attacks. In this paper, we present a novel framework, i.e., Transferable Hypergraph Attack via Injecting Nodes into Pivotal Hyperedges (TH-Attack), to address these limitations. Specifically, we design a hyperedge recognizer via pivotality assessment to obtain pivotal hyperedges within the aggregation paths of HGNNs. Furthermore, we introduce a feature inverter based on pivotal hyperedges, which generates malicious nodes by maximizing the semantic divergence between the generated features and the pivotal hyperedges features. Lastly, by injecting these malicious nodes into the pivotal hyperedges, TH-Attack improves the transferability and effectiveness of attacks. Extensive experiments are conducted on six authentic datasets to validate the effectiveness of TH-Attack and the corresponding superiority to state-of-the-art methods. Meixia He, Peican Zhu, Yangming Guo, Manman Yuan, Keke Tang |
AAAI | 6 |
| 2026 | Less Is More: Sparse and Cooperative Perturbation for Point Cloud AttacksabstractMost adversarial attacks on point clouds perturb a large number of points, causing widespread geometric changes and limiting applicability in real-world scenarios. While recent works explore sparse attacks by modifying only a few points, such approaches often struggle to maintain effectiveness due to the limited influence of individual perturbations. In this paper, we propose SCP, a sparse and cooperative perturbation framework that selects and leverages a compact subset of points whose joint perturbations produce amplified adversarial effects. Specifically, SCP identifies the subset where the misclassification loss is locally convex with respect to their joint perturbations, determined by checking the positive-definiteness of the corresponding Hessian block. The selected subset is then optimized to generate high-impact adversarial examples with minimal modifications. Extensive experiments show that SCP achieves 100% attack success rates, surpassing state-of-the-art sparse attacks, and delivers superior imperceptibility to dense attacks with far fewer modifications. Keke Tang, Tianyu Hao, Weilong Peng, Denghui Zhang 0001, Peican Zhu, Zhihong Tian 0001 |
AAAI | 1 |
| 2026 | WaveSculpt: Text-to-3D generation with wavelet-guided score distillation
Weilong Peng, Jianhui Huo, Keke Tang, Yangtao Wang, Yan Wang 0022, Meie Fang |
Comput. Aided Geom. Des. | 3 |
| 2026 | Transferable and undefendable point cloud attacks via medial axis transform
Keke Tang, Yuze Gao, Weilong Peng, Meie Fang, Peican Zhu |
Comput. Aided Geom. Des. | 1 |
| 2026 | Transferable Black-Box Injection Attack Against Heterogeneous Graph Neural NetworksabstractRecent studies have shown that Heterogeneous Graph Neural Networks (HetGNNs) are vulnerable to adversarial attacks. Existing methods rely on the gradient information of source models or surrogate models to generate perturbations, which limits the attack of transferability and effectiveness in real-world attack scenarios, while consuming time costs. In this paper, we propose a novel framework, i.e., Transferable Black-Box Injection Attack against Heterogeneous Graph Neural Networks (TBI-Attack), to address these challenges. Specifically, we introduce a voting-based key nodes recognizer based on different meta-paths to identify key nodes in relational subgraphs. Subsequently, we present a semantic-confused feature generator that leverages self-supervised learning to integrate neighborhood information from different relational subgraphs, generating malicious nodes with conflicting semantic features. Furthermore, malicious nodes are injected into various relational subgraphs, disrupting their specific semantic functionality and systematically impairing the message-passing process of HetGNNs. Extensive experiments are conducted on four authentic datasets to validate the effectiveness and transferability of TBI-Attack and the corresponding superiority to state-of-the-art methods. Meixia He, Peican Zhu, Jianrui Chen 0002, Keke Tang, Zhen Wang 0004 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Hard-Label Black-Box Attacks on 3D Point CloudsabstractWith the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial attacks. Almost all existing 3D attackers simply follow the white-box or black-box setting to iteratively update coordinate perturbations based on back-propagated or estimated gradients. However, these methods are hard to deploy in real-world scenarios (no model details are provided) as they severely rely on parameters or output logits of victim models. To this end, we propose point cloud attacks from a more practical setting,i.e., hard-label black-box attack, in which attackers can only access the prediction label of 3D input. We introduce a novel 3D attack method based on a new spectrum-aware decision boundary algorithm to generate high-quality adversarial samples. In particular, we first construct a class-aware model decision boundary, by developing a learnable spectrum-fusion strategy to adaptively fuse point clouds of different classes in the spectral domain, aiming to craft their intermediate samples without distorting the original geometry. Then, we devise an iterative coordinate-spectrum optimization method with curvature-aware boundary search to move the intermediate sample along the decision boundary for generating adversarial point clouds with trivial perturbations. Experiments demonstrate that our attack competitively outperforms existing white/black-box attackers in terms of attack performance and adversary quality. Daizong Liu, Yunbo Tao, Junhao Dong 0001, Keke Tang, Pan Zhou 0001, Wei Hu 0003, Yew-Soon Ong |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkabstractGiven some video-query pairs with untrimmed videos and sentence queries, temporal sentence grounding (TSG) aims to locate query-relevant segments in these videos. Although previous respectable TSG methods have achieved remarkable success, they train each video-query pair separately and ignore the relationship between different pairs. To this end, in this paper, we pose a brand-new setting: Multi-Pair TSG, which aims to co-train these pairs. We propose a novel video-query co-training approach, Multi-Thread Knowledge Transfer Network, to locate a variety of video-query pairs effectively and efficiently. Firstly, we mine the spatial and temporal semantics across different queries to cooperate with each other. To learn intra- and inter-modal representations simultaneously, we design a cross-modal contrast module to explore the semantic consistency by a self-supervised strategy. To fully align visual and textual representations between different pairs, we design a prototype alignment strategy to 1) match object prototypes and phrase prototypes for spatial alignment, and 2) align activity prototypes and sentence prototypes for temporal alignment. Finally, we develop an adaptive negative selection module to adaptively generate a threshold for cross-modal matching. Extensive experiments show the effectiveness and efficiency of our proposed method. Wanlong Fang, Changshuo Wang 0001, Daizong Liu, Keke Tang, Jianfeng Dong, Pan Zhou 0001, Beibei Li 0002 |
AAAI | 5 |
| 2025 | Hypergraph Attacks via Injecting Homogeneous Nodes into Elite HyperedgesabstractRecent studies have shown that Hypergraph Neural Networks (HGNNs) are vulnerable to adversarial attacks. Existing approaches focus on hypergraph modification attacks guided by gradients, overlooking node spanning in the hypergraph and the group identity of hyperedges, thereby resulting in limited attack performance and detectable attacks. In this manuscript, we present a novel framework, i.e., Hypergraph Attacks via Injecting Homogeneous Nodes into Elite Hyperedges (IE-Attack), to tackle these challenges. Initially, utilizing the node spanning in the hypergraph, we propose the elite hyperedges sampler to identify hyperedges to be injected. Subsequently, a node generator utilizing Kernel Density Estimation (KDE) is proposed to generate the homogeneous node with the group identity of hyperedges. Finally, by injecting the homogeneous node into elite hyperedges, IE-Attack improves the attack performance and enhances the imperceptibility of attacks. Extensive experiments are conducted on five authentic datasets to validate the effectiveness of IE-Attack and the corresponding superiority to state-of-the-art methods. Meixia He, Peican Zhu, Keke Tang, Yangming Guo |
AAAI | 3 |
| 2025 | Imperceptible 3D Point Cloud Attacks on Lattice-based Barycentric CoordinatesabstractImperceptible adversarial attacks on 3D point clouds rely on effective constraints. While manifold constraints have notable advantages over Euclidean ones, the global parameterization used in current methods often fails to fully preserve manifold properties. In this paper, we propose to constrain lattice-based barycentric coordinates during attacks from a local parametric perspective to ensure imperceptibility. Specifically, we utilize a permutohedral lattice to partition point clouds into multiple cells, and then extract barycentric coordinates for each point within these cells, forming a local parametric representation of the point clouds. By enforcing local parametric constraints that minimize the displacement of barycentric coordinates, we largely preserve the manifold properties, ultimately leading to improved imperceptibility. Extensive experiments validate that integrating these local parametric constraints into conventional adversarial attacks yields superior imperceptibility, outperforming state-of-the-art methods. Keke Tang, Ziyong Du, Weilong Peng, Daizong Liu, Ligang Liu 0001, Zhihong Tian 0001 |
AAAI | 1 |
| 2025 | EvoAgents: A Cognitive-Driven Framework for Personality Evolution in Generative Agent Society
Keke Tang, Cui Tang |
CogSci | 5 |
| 2025 | Training-Free Language-Guided Video Summarization via Multi-Grained Saliency Scoring
Yongwei Nie, Fei Ma 0006, Keke Tang, F. Richard Yu, Hongmin Cai, Ping Li 0016 |
CVM (3) | 4 |
| 2025 | Simplification Is All You Need against Out-of-Distribution OverconfidenceabstractDeep neural networks (DNNs) often exhibit out-of-distribution (OOD) overconfidence, producing overly confident predictions on OOD samples. We attribute this issue to the inherent over-complexity of DNNs and investigate two key aspects: capacity and nonlinearity. First, we demonstrate that reducing model capacity through knowledge distillation can effectively mitigate OOD overconfidence. Second, we show that selectively reducing nonlinearity by removing ReLU operations further alleviates the issue. Building on these findings, we present a practical guide to model simplification, combining both strategies to significantly reduce OOD overconfidence. Extensive experiments validate the effectiveness of this approach in mitigating OOD overconfidence and demonstrate its superiority over state-of-the-art methods. Additionally, our simplification strategies can be combined with existing OOD detection techniques to further enhance OOD detection performance. Keke Tang, Weilong Peng, Zhize Wu, Yongwei Nie, Wenping Wang 0001, Zhihong Tian 0001 |
CVPR | 1 |
| 2025 | Imperceptible Adversarial Attacks on Point Clouds Guided by Point-to-Surface FieldabstractAdversarial attacks on point clouds are crucial for assessing and improving the adversarial robustness of 3D deep learning models. Traditional solutions strictly limit point displacement during attacks, making it challenging to balance imperceptibility with adversarial effectiveness. In this paper, we attribute the inadequate imperceptibility of adversarial attacks on point clouds to deviations from the underlying surface. To address this, we introduce a novel point-to-surface (P2S) field that adjusts adversarial perturbation directions by dragging points back to their original underlying surface. Specifically, we use a denoising network to learn the gradient field of the logarithmic density function encoding the shape’s surface, and apply a distance-aware adjustment to perturbation directions during attacks, thereby enhancing imperceptibility. Extensive experiments show that adversarial attacks guided by our P2S field are more imperceptible, outperforming state-of-the-art methods. Keke Tang, Weiyao Ke, Weilong Peng, Ziyong Du, Zhize Wu, Peican Zhu, Zhihong Tian 0001 |
ICASSP | 1 |
| 2025 | Attribute-Based Out-of-Distribution Detection Using LLaVA
Daojie Zhao, Yongwei Nie, Peican Zhu, Keke Tang |
ICIC (19) | 5 |
| 2025 | Imperceptible Beam-Sensitive Adversarial Attacks for LiDAR-based Object Detection in Autonomous DrivingabstractLiDAR-based 3D perception plays a pivotal role in autonomous driving systems, which is a crucial component facilitating obstacle avoidance for vehicles, posing potential security risks in daily usage. Previous LiDAR-based adversarial attacks against autonomous driving models generally focus on the elimination or modification of specific objects in the point cloud. However, such attacks severely rely on the prior knowledge of the location of the target object and can only have limited adversarial impacts within certain local detection results. Instead, in this paper, we develop a novel LiDAR-based 3D adversarial attack from a more comprehensive perspective, which adaptively learns to impose global hazards without choosing particular local point sets of genuine obstacles, while enhancing the effectiveness, imperceptibility, and practicality of the perturbed LiDAR point clouds. We conduct extensive experiments on the large self-driving dataset Waymo to demonstrate the effectiveness of our attack against four object detection models (CenterPoint, PV-RCNN, DSVT, VoxelNext). Fuyao Cai, Daizong Liu, Jixiang Yu, Keke Tang, Pan Zhou 0001 |
ICME | 5 |
| 2025 | PopuDet: Autism Spectrum Disorder Detection in Population Graphs via Micro-macro Relationship Construction and Multi-feature FusionabstractPopulation graphs are crucial for assessing clinical risk and enhancing the accuracy of Autism Spectrum Disorder (ASD) detection. Nevertheless, the current population graph construction overlooks the balance between biological signals and clinical manifestations, leading to relationship deviation within the population graph and poor detection performance. To address this challenge, we propose a novel approach for ASD Detection in Population Graphs (PopuDet) via Micro-macro Relationship Construction (MmRC) and Multi-feature Fusion (MF). Specifically, our method utilizes the MmRC module to construct a multi-scale population graph balancing the relationships between biological signals synchrony and clinical subtype groups. Subsequently, the MF module learns high-level graph representations at different scales for adaptive fusion to achieve precise ASD detection. Extensive experiments validate the efficiency of PopuDet, highlighting its superior performance over current state-of-the-art methods. Our source code is available at https://github.com/xuting99/PopuDet. Manman Yuan, Jiazhen Ye, Peican Zhu, Keke Tang |
ICME | 6 |
| 2025 | SourceDetMamba: A Graph-aware State Space Model for Source Detection in Sequential HypergraphsabstractSource detection on graphs has demonstrated high efficacy in identifying rumor origins. Despite advances in machine learning-based methods, many fail to capture intrinsic dynamics of rumor propagation. In this work, we present SourceDetMamba: A Graph-aware State Space Model for Source Detection in Sequential Hypergraphs, which harnesses the recent success of the state space model Mamba, known for its superior global modeling capabilities and computational efficiency, to address this challenge. Specifically, we first employ hypergraphs to model high-order interactions within social networks. Subsequently, temporal network snapshots generated during the propagation process are sequentially fed in reverse order into Mamba to infer underlying propagation dynamics. Finally, to empower the sequential model to effectively capture propagation patterns while integrating structural information, we propose a novel graph-aware state update mechanism, wherein the state of each node is propagated and refined by both temporal dependencies and topological context. Extensive evaluations on eight datasets demonstrate that SourceDetMamba consistently outperforms state-of-the-art approaches. Peican Zhu, Yangming Guo, Chao Gao 0001, Zhen Wang 0004, Keke Tang |
IJCAI | 6 |
| 2025 | HyperDet: Source Detection in Hypergraphs via Interactive Relationship Construction and Feature-rich Attention FusionabstractHypergraphs offer superior modeling capabilities for social networks, particularly in capturing group phenomena that extend beyond pairwise interactions in rumor propagation. Existing approaches in rumor source detection predominantly focus on dyadic interactions, which inadequately address the complexity of more intricate relational structures. In this study, we present a novel approach for Source Detection in Hypergraphs (HyperDet) via Interactive Relationship Construction and Feature-rich Attention Fusion. Specifically, our methodology employs an Interactive Relationship Construction module to accurately model both the static topology and dynamic interactions among users, followed by the Feature-rich Attention Fusion module, which autonomously learns node features and discriminates between nodes using a self-attention mechanism, thereby effectively learning node representations under the framework of accurately modeled higher-order relationships. Extensive experimental validation confirms the efficacy of our HyperDet approach, showcasing its superiority relative to current state-of-the-art methods. Peican Zhu, Yangming Guo, Keke Tang, Chao Gao 0001, Zhen Wang 0004 |
IJCAI | 4 |
| 2025 | AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical PerspectiveabstractAdversarial attacks on robotic grasping provide valuable insights into evaluating and improving the robustness of these systems. Unlike studies that focus solely on neural network predictions while overlooking the physical principles of grasping, this paper introduces AdvGrasp, a framework for adversarial attacks on robotic grasping from a physical perspective. Specifically, AdvGrasp targets two core aspects: lift capability, which evaluates the ability to lift objects against gravity, and grasp stability, which assesses resistance to external disturbances. By deforming the object's shape to increase gravitational torque and reduce stability margin in the wrench space, our method systematically degrades these two key grasping metrics, generating adversarial objects that compromise grasp performance. Extensive experiments across diverse scenarios validate the effectiveness of AdvGrasp, while real-world validations demonstrate its robustness and practical applicability. Mingliang Han, Tianyu Hao, Cegang Li, Yun-Bo Zhao, Keke Tang |
IJCAI | 6 |
| 2025 | EOOD: Entropy-based Out-of-distribution DetectionabstractDeep neural networks (DNNs) often exhibit overconfidence when encountering out-of-distribution (OOD) samples, posing significant challenges for deployment. Since DNNs are trained on in-distribution (ID) datasets, the information flow of ID samples through DNNs inevitably differs from that of OOD samples. In this paper, we propose an Entropy-based Out-Of-distribution Detection (EOOD) framework. EOOD first identifies specific block where the information flow differences between ID and OOD samples are more pronounced, using both ID and pseudo-OOD samples. It then calculates the conditional entropy on the selected block as the OOD confidence score. Comprehensive experiments conducted across various ID and OOD settings demonstrate the effectiveness of EOOD in OOD detection and its superiority over state-of-the-art methods. Guide Yang, Weilong Peng, Yongwei Nie, Peican Zhu, Keke Tang |
IJCNN | 7 |
| 2025 | KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News DetectionabstractIn recent years, the rampant spread of misinformation on social media has made accurate detection of multimodal fake news a critical research focus. However, previous research has not adequately understood the semantics of images, and models struggle to discern news authenticity with limited textual information. Meanwhile, treating all emotional types of news uniformly without tailored approaches further leads to performance degradation. Therefore, we propose a novel Knowledge Augmentation and Emotion Guidance Network (KEN). On the one hand, we effectively leverage LVLM's powerful semantic understanding and extensive world knowledge. For images, the generated captions provide a comprehensive understanding of image content and scenes, while for text, the retrieved evidence helps break the information silos caused by the closed and limited text and context. On the other hand, we consider inter-class differences between different emotional types of news through balanced learning, achieving fine-grained modeling of the relationship between emotional types and authenticity. Extensive experiments on two real-world datasets demonstrate the superiority of our KEN. Peican Zhu, Yubo Jing, Keke Tang, Yangming Guo |
ACM Multimedia | 4 |
| 2025 | EIA: Edge-Aware Imperceptible Adversarial Attacks on 3D Point Clouds
Zhensu Wang, Weilong Peng, Le Wang 0008, Zhizhe Wu, Peican Zhu, Keke Tang |
MMM (1) | 6 |
| 2025 | Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsabstractAlthough Large Vision-Language Models (LVLMs) exhibit impressive multimodal capabilities, their vulnerability to adversarial examples has raised serious security concerns.
Existing LVLM attackers simply optimize adversarial images that easily overfit a certain model/prompt, making them ineffective once they are transferred to attack a different model/prompt.
Motivated by this research gap, this paper aims to develop a more powerful attack that is transferable to black-box LVLM models of different structures and task-aware prompts of different semantics.
Specifically, we introduce a new perspective of information theory to investigate LVLMs' transferable characteristics by exploring the relative dependence between outputs of the LVLM model and input adversarial samples. Our empirical observations suggest that enlarging/decreasing the mutual information between outputs and the disentangled adversarial/benign patterns of input images helps to generate more agnostic perturbations for misleading LVLMs' perception with better transferability.
In particular, we formulate the complicated calculation of information gain as an estimation problem and incorporate such informative constraints into the adversarial learning process.
Extensive experiments on various LVLM models/prompts demonstrate our significant transfer-attack performance. Xiaowen Cai 0001, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Keke Tang, Pan Zhou 0001, Lichao Sun 0001, Wei Hu 0003 |
NeurIPS | 6 |
| 2025 | SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News DetectionabstractPrevious studies on multimodal fake news detection mainly focus on the alignment and integration of cross-modal features, as well as the application of text-image consistency. However, they overlook the semantic enhancement effects of large multimodal models and pay little attention to the emotional features of news. In addition, people find that fake news is more inclined to contain negative emotions than real ones. Therefore, we propose a novel Semantic Enhancement and Emotional Reasoning (SEER) Network for multimodal fake news detection. We generate summarized captions for image semantic understanding and utilize the products of large multimodal models for semantic enhancement. Inspired by the perceived relationship between news authenticity and emotional tendencies, we propose an expert emotional reasoning module that simulates real-life scenarios to optimize emotional features and infer the authenticity of news. Extensive experiments on two real-world datasets demonstrate the superiority of our SEER over state-of-the-art baselines. Peican Zhu, Yubo Jing, Lianwei Wu, Keke Tang |
SMC | 7 |
| 2025 | MeshPAD: Payload-aware mesh distortion for 3D steganography based on geometric deep learning
Weilong Peng, Keke Tang, Weixuan Tang 0002, Yong Su 0003, Meie Fang, Ping Li 0016 |
Expert Syst. Appl. | 2 |
| 2025 | Efficient Source Detection in Incomplete Networks via Sensor Deployment and Source ApproachingabstractRumor source detection in structurally incomplete networks holds significant practical importance. Existing methods predominantly assume a complete network structure information; furthermore, they often neglect the issue of resource consumption, i.e., sensor deployment. In this paper, we propose an efficient source detection approach in incomplete networks via propagation-aware Sensor Deployment and time stamp-guided Source Approaching (SDSA) to tackle these challenges. Specifically, during the sensor deployment phase, we employ quality-guaranteed Monte Carlo propagation simulations coupled with a greedy strategy to achieve maximum coverage with minimal sensors. In the source detection phase, for the structurally incomplete network snapshots, we first attempt edge reconnection from the sensor with the earliest timestamp, followed by posterior maximization Bayesian estimation for source identification. Extensive experiments demonstrate the effectiveness of SDSA and its superiority over state-of-the-art methods. The code has been made publicly available athttps://github.com/cheng-le/SDSA. Peican Zhu, Keke Tang, Chao Gao 0001, Zhen Wang 0004 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Continuous Bijection Supervised Pyramid Diffeomorphic Deformation for Learning Tooth Meshes From CBCT ImagesabstractAccurate and high-quality tooth mesh generation from cone-beam computerized tomography (CBCT) is an essential computer-aided technology for digital dentistry. However, existing segmentation-based methods require complicated post-processing and significant manual correction to generate regular tooth meshes. In this paper, we propose a method of continuous bijection supervised pyramid diffeomorphic deformation (PDD) for learning tooth meshes, which could be used to directly generate high-quality tooth meshes from CBCT Images. Overall, we adopt a classic two-stage framework. In the first stage, we devise an enhanced detector to accurately locate and crop every tooth. In the second stage, a PDD network is designed to deform a sphere mesh from low resolution to high one according to pyramid flows based on diffeomorphic mesh deformations, so that the generated mesh approximates the ground truth infinitely and efficiently. To achieve that, a novel continuous bijection distance loss on the diffeomorphic sphere is also designed to supervise the deformation learning, which overcomes the shortcoming of loss based on nearest-neighbour mapping and improves the fitting precision. Experiments show that our method outperforms the state-of-the-art methods in terms of both different evaluation metrics and the geometry quality of reconstructed tooth surfaces. Zechu Zhang, Weilong Peng, Jinyu Wen, Keke Tang, Meie Fang, David Dagan Feng, Ping Li 0016 |
IEEE Trans. Multim. | 4 |
| 2025 | Decision Fusion Networks for Image ClassificationabstractConvolutional neural networks, in which each layer receives features from the previous layer(s) and then aggregates/abstracts higher level features from them, are widely adopted for image classification. To avoid information loss during feature aggregation/abstraction and fully utilize lower layer features, we propose a novel decision fusion module (DFM) for making an intermediate decision based on the features in the current layer and then fuse its results with the original features before passing them to the next layers. This decision is devised to determine an auxiliary category corresponding to the category at a higher hierarchical level, which can, thus, serve as category-coherent guidance for later layers. Therefore, by stacking a collection of DFMs into a classification network, the generated decision fusion network is explicitly formulated to progressively aggregate/abstract more discriminative features guided by these decisions and then refine the decisions based on the newly generated features in a layer-by-layer manner. Comprehensive results on four benchmarks validate that the proposed DFM can bring significant improvements for various common classification networks at a minimal additional computational cost and are superior to the state-of-the-art decision fusion-based methods. In addition, we demonstrate the generalization ability of the DFM to object detection and semantic segmentation. Keke Tang, Yuexin Ma, Dingruibo Miao, Peng Song 0001, Zhaoquan Gu, Zhihong Tian 0001, Wenping Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | GIN-SD: Source Detection in Graphs with Incomplete Nodes via Positional Encoding and Attentive FusionabstractSource detection in graphs has demonstrated robust efficacy in the domain of rumor source identification. Although recent solutions have enhanced performance by leveraging deep neural networks, they often require complete user data. In this paper, we address a more challenging task, rumor source detection with incomplete user data, and propose a novel framework, i.e., Source Detection in Graphs with Incomplete Nodes via Positional Encoding and Attentive Fusion (GIN-SD), to tackle this challenge. Specifically, our approach utilizes a positional embedding module to distinguish nodes that are incomplete and employs a self-attention mechanism to focus on nodes with greater information transmission capacity. To mitigate the prediction bias caused by the significant disparity between the numbers of source and non-source nodes, we also introduce a class-balancing mechanism. Extensive experiments validate the effectiveness of GIN-SD and its superiority to state-of-the-art methods. Peican Zhu, Keke Tang, Chao Gao 0001, Zhen Wang 0004 |
AAAI | 3 |
| 2024 | Manifold Constraints for Imperceptible Adversarial Attacks on Point CloudsabstractAdversarial attacks on 3D point clouds often exhibit unsatisfactory imperceptibility, which primarily stems from the disregard for manifold-aware distortion, i.e., distortion of the underlying 2-manifold surfaces. In this paper, we develop novel manifold constraints to reduce such distortion, aiming to enhance the imperceptibility of adversarial attacks on 3D point clouds. Specifically, we construct a bijective manifold mapping between point clouds and a simple parameter shape using an invertible auto-encoder. Consequently, manifold-aware distortion during attacks can be captured within the parameter space. By enforcing manifold constraints that preserve local properties of the parameter shape, manifold-aware distortion is effectively mitigated, ultimately leading to enhanced imperceptibility. Extensive experiments demonstrate that integrating manifold constraints into conventional adversarial attack solutions yields superior imperceptibility, outperforming the state-of-the-art methods. Keke Tang, Weilong Peng, Jianpeng Wu, Yawen Shi, Daizong Liu, Pan Zhou 0001, Wenping Wang 0001, Zhihong Tian 0001 |
AAAI | 1 |
| 2024 | Towards Robust Temporal Activity Localization Learning with Noisy LabelsabstractThis paper addresses the task of temporal activity localization (TAL). Although recent works have made significant progress in TAL research, almost all of them implicitly assume that the dense frame-level correspondences in each video-query pair are correctly annotated. However, in reality, such an assumption is extremely expensive and even impossible to satisfy due to subjective labeling. To alleviate this issue, in this paper, we explore a new TAL setting termed Noisy Temporal activity localization (NTAL), where a TAL model should be robust to the mixed training data with noisy moment boundaries. Inspired by the memorization effect of neural networks, we propose a novel method called Co-Teaching Regularizer (CTR) for NTAL. Specifically, we first learn a Gaussian Mixture Model to divide the mixed training data into preliminary clean and noisy subsets. Subsequently, we refine the labels of the two subsets by an adaptive prediction function so that their true positive and false positive samples could be identified. To avoid single model being prone to its mistakes learned by the mixed data, we adopt a co-teaching paradigm, which utilizes two models sharing the same framework to teach each other for robust learning. A curriculum strategy is further introduced to gradually learn the moment confidence from easy to hard. Experiments on three datasets demonstrate that our CTR is significantly more robust to the noisy training data compared to the existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Guoshun Nan, Keke Tang, Wanlong Fang, Yu Cheng 0001 |
LREC/COLING | 7 |
| 2024 | CORES: Convolutional Response-based Score for Out-of-distribution DetectionabstractDeep neural networks (DNNs) often display overconfidence when encountering out-of-distribution (OOD) samples, posing significant challenges in real-world applications. Capitalizing on the observation that responses on convolutional kernels are generally more pronounced for in-distribution (ID) samples than for OOD ones, this paper proposes the COnvolutional REsponse-based Score (CORES) to exploit these discrepancies for OOD detection. Initially, CORES delves into the extremities of convolutional responses by considering both their magnitude and the frequency of significant values. Moreover, through backtracking from the most prominent predictions, CORES effectively pinpoints sample-relevant kernels across different layers. These kernels, which exhibit a strong correlation to input samples, are integral to CORES's OOD detection capability. Comprehensive experiments across various ID and OOD settings demonstrate CORES's effectiveness in OOD detection and its superiority to the state-of-the-art methods. Keke Tang, Weilong Peng, Runnan Chen, Peican Zhu, Wenping Wang 0001, Zhihong Tian 0001 |
CVPR | 1 |
| 2024 | Rethinking Weakly-Supervised Video Temporal Grounding From a Game Perspective
Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen 0006, Jianfeng Dong, Keke Tang, Pan Zhou 0001, Yu Cheng 0001, Daizong Liu |
ECCV (45) | 7 |
| 2024 | FLAT: Flux-Aware Imperceptible Adversarial Attacks on 3D Point Clouds
Keke Tang, Lujie Huang, Weilong Peng, Daizong Liu, Ligang Liu 0001, Zhihong Tian 0001 |
ECCV (6) | 1 |
| 2024 | Hiding Imperceptible Noise in Curvature-Aware Patches for 3D Point Cloud Attack
Daizong Liu, Keke Tang, Pan Zhou 0001, Lixing Chen, Junyang Chen 0001 |
ECCV (30) | 3 |
| 2024 | GAA-BD: Graph Adversarial Augmentation-based Social Bot DetectionabstractOnline social networks are crucial for information acquisition nowadays, yet they are increasingly jeopardized by malicious attacks from social bots. This underscores the urgent need for robust social bot detection methods. To address the significant challenge posed by the imbalance in the number of human and bot users, we introduce a Graph Adversarial Augmentation-based method for social Bot Detection (GAA-BD) to improve detection efficacy. Our method incorporates a graph convolutional neural network as its core architecture and enhances it with adversarial augmentation applied to both the dataset and the training process. By strategically generating synthetic samples and employing targeted adversarial training, our approach effectively resolves sample size imbalances and bolsters model robustness. Comprehensive experimental results demonstrate that our proposed method outperforms existing baselines, establishing its effectiveness in combatting social bot infiltration in online networks. Keke Tang, Peican Zhu |
HPCC | 5 |
| 2024 | Reparameterization Head for Efficient Multi-Input NetworksabstractReparameterization techniques have demonstrated their efficacy in improving the efficiency of deep neural networks. However, their application has been largely confined to single-input network structures, leaving multi-input ones, commonly encountered in real-world applications, largely unexplored. In this paper, we formulate reparameterization head (RepHead), the first framework designed to introduce reparameterization into multi-input neural networks. RepHead compresses multiple inputs into a single input and employs reconstruction operations to recover them, thereby transforming multi-input networks into single-input, multibranch architectures, thereby enabling the application of reparameterization. We demonstrate the usage of RepHead in both image and point cloud domains. Extensive experimental results validate that the integration of RepHead substantially reduces computational overhead and memory requirements while maintaining minimal performance loss. Keke Tang, Weilong Peng, Peican Zhu, Zhihong Tian 0001 |
ICASSP | 1 |
| 2024 | IE-aware Consistency Losses for Detailed 3D Face Reconstruction from Multiple Images in the Wildabstract3D face reconstruction from multiple in-the-wild images in an unsupervised manner poses a significant challenge, primarily due to the pervasive presence of Intrinsic and Extrinsic inconsistencies in facial features. To tackle this, we introduce a novel set of IE-aware consistency losses designed to effectively mitigate these inconsistencies. Our Local Alignment Loss employs neighborhood search techniques to identify and optimize consistent pixel information, thereby reducing intrinsic inconsistencies. In parallel, our Region Subset Selection Loss filters out regions where significant discrepancies exist between the input and reconstructed images, effectively alleviating extrinsic inconsistencies. Extensive experimental results validate the effectiveness of our IE-aware consistency losses in reconstructing detailed 3D facial geometry from images captured in uncontrolled environments. Weilong Peng, Keke Tang, Kongyang Chen, Yangtao Wang, Ping Li 0016, Meie Fang |
ICME | 3 |
| 2024 | Leveraging LLMs to Enhance NLP Performance Through Distillation and Optimized Training Strategies
Yining Huang, Keke Tang, Wanmin Lian, Meilian Chen |
ICONIP (10) | 2 |
| 2024 | A General Black-box Adversarial Attack on Graph-based Fake News Detectors
Peican Zhu, Zechen Pan, Yang Liu 0144, Jiwei Tian, Keke Tang, Zhen Wang 0004 |
IJCAI | 5 |
| 2024 | Frequency-Aware GAN for Imperceptible Transfer Attack on 3D Point CloudsabstractWith the development of depth sensors and 3D vision, the vulnerability of 3D point cloud models has garnered heightened concern. Almost all existing 3D attackers are deployed in the white-box setting, where they access the model details and directly optimize coordinate-wise noises to perturb 3D objects. However, realistic 3D applications would not share any model information (model parameters, gradients, etc.) with users. Although a few recent works try to explore the black-box attack, they still achieve limited attack success rates (ASR) and fail to generate high-quality adversarial samples. In this paper, we focus on designing a transfer-based black-box attack method, called Transferable Frequency-aware 3D GAN, to delve into achieving a high black-box ASR by improving the adversarial transferability while making the adversarial samples more imperceptible. Considering that the 3D imperceptibility depends on whether the shape of the object is distorted, we utilize the spectral tool with the GAN design to explicitly perceive and preserve the 3D geometric structures. Specifically, we design the Graph Fourier Transform (GFT) encoding layer in the GAN generator to extract the geometries as guidance, and develop a corresponding Inverse-GFT decoding layer to decode latent features with this guidance to reconstruct high-quality adversarial samples. To further improve the transferability, we develop a dual learning scheme of discriminator from both frequency and feature perspectives to constrain the generator via adversarial learning. Finally, imperceptible and transferable perturbations are rapidly generated by our proposed attack. Experimental results demonstrate that our attack method achieves the highest transfer ASR while exhibiting stronger imperceptibility. Xiaowen Cai 0001, Yunbo Tao, Daizong Liu, Pan Zhou 0001, Xiaoye Qu, Jianfeng Dong, Keke Tang, Lichao Sun 0001 |
ACM Multimedia | 7 |
| 2024 | SymAttack: Symmetry-aware Imperceptible Adversarial Attacks on 3D Point CloudsabstractAdversarial attacks on point clouds are crucial for assessing and improving the adversarial robustness of 3D deep learning models. Despite leveraging various geometric constraints, current adversarial attack strategies often suffer from inadequate imperceptibility. Given that adversarial perturbations tend to disrupt the inherent symmetry in objects, we recognize this disruption as the primary cause of the lack of imperceptibility in these attacks. In this paper, we introduce a novel framework, symmetry-aware imperceptible adversarial attacks on 3D point clouds (SymAttack), to address this issue. Our approach starts by identifying part- and patch-level symmetry elements, and grouping points based on semantic and Euclidean distances, respectively. During the adversarial attack iterations, we intentionally adjust the perturbation vectors on symmetric points relative to their symmetry plane. By preserving symmetry within the attack process, SymAttack significantly enhances imperceptibility. Extensive experiments validate the effectiveness of SymAttack in generating imperceptible adversarial point clouds, demonstrating its superiority over the state-of-the-art methods. Keke Tang, Zhensu Wang, Weilong Peng, Lujie Huang, Le Wang 0008, Peican Zhu, Wenping Wang 0001, Zhihong Tian 0001 |
ACM Multimedia | 1 |
| 2024 | Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding tasks. Nevertheless, these models are susceptible to adversarial examples. In real-world applications, existing LVLM attackers generally rely on the detailed prior knowledge of the model to generate effective perturbations. Moreover, these attacks are task-specific, leading to significant costs for designing perturbation. Motivated by the research gap and practical demands, in this paper, we make the first attempt to build a universal attacker against real-world LVLMs, focusing on two critical aspects: (i) restricting access to only the LVLM inputs and outputs. (ii) devising a universal adversarial patch, which is task-agnostic and can deceive any LVLM-driven task when applied to various inputs. Specifically, we start by initializing the location and the pattern of the adversarial patch through random sampling, guided by the semantic distance between their output and the target label. Subsequently, we maintain a consistent patch location while refining the pattern to enhance semantic resemblance to the target. In particular, our approach incorporates a diverse set of LVLM task inputs as query samples to approximate the patch gradient, capitalizing on the importance of distinct inputs. In this way, the optimized patch is universally adversarial against different tasks and prompts, leveraging solely gradient estimates queried from the model. Extensive experiments are conducted to verify the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs including LLaVA, MiniGPT-4, Flamingo, and BLIP-2, spanning a spectrum of tasks, all achieved without delving into the details of the model structures. Daizong Liu, Xiaoye Qu, Pan Zhou 0001, Keke Tang, Yao Wan 0001, Lichao Sun 0001 |
NeurIPS | 6 |
| 2024 | LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed EffectabstractLanguage models have shown vulnerability against backdoor attacks, threatening the security of services based on them. To mitigate the threat, existing solutions attempted to search for backdoor triggers, which can be time-consuming when handling a large search space. Looking into the attack process, we observe that poisoned data will create a long-tailed effect in the victim model, causing the decision boundary to shift towards the attack targets. Inspired by this observation, we introduce LT-Defense, the first searching-free backdoor defense via exploiting the long-tailed effect. Specifically, LT-Defense employs a small set of clean examples and two metrics to distinguish backdoor-related features in the target model. Upon detecting a backdoor model, LT-Defense additionally provides test-time backdoor freezing and attack target prediction. Extensive experiments demonstrate the effectiveness of LT-Defense in both detection accuracy and efficiency, e.g., in task-agnostic scenarios, LT-Defense achieves 98% accuracy across 1440 models with less than 1% of the time cost of state-of-the-art solutions. Yixiao Xu, Binxing Fang, Mohan Li, Keke Tang, Zhihong Tian 0001 |
NeurIPS | 4 |
| 2024 | MIT: Multi-cue Injected Transformer for Two-Stage HOI Detection
Weilong Peng, Qingfeng Chen, Keke Tang, Meng Xing, Meie Fang |
PRCV (7) | 3 |
| 2024 | Improving adversarial transferability through hybrid augmentation
Peican Zhu, Zepeng Fan, Sensen Guo, Keke Tang |
Comput. Secur. | 4 |
| 2024 | A knowledge-guided graph attention network for emotion-cause pair extraction
Peican Zhu, Keke Tang, Haifeng Zhang 0003, Zhen Wang 0004 |
Knowl. Based Syst. | 3 |
| 2024 | Node Injection Attack Based on Label Propagation Against Graph Neural NetworkabstractGraph neural network (GNN) has achieved remarkable success in various graph learning tasks, such as node classification, link prediction, and graph classification. The key to the success of GNN lies in its effective structure information representation through neighboring aggregation. However, the attacker can easily perturb the aggregation process through injecting fake nodes, which reveals that GNN is vulnerable to the graph injection attack (GIA). Existing GIA methods primarily focus on damaging the classical feature aggregation process while overlooking the neighborhood aggregation process via label propagation. To bridge this gap, we propose the label-propagation-based global injection attack (LPGIA) which conducts the GIA on the node classification task. Specifically, we analyze the aggregation process from the perspective of label propagation and transform the GIA problem into a global injection label specificity attack problem. To solve this problem, LPGIA utilizes a label-propagation-based strategy to optimize the combinations of the nodes connected to the injected node. Then, LPGIA leverages the feature mapping to generate malicious features for injected nodes. In extensive experiments against representative GNNs, LPGIA outperforms the previous best-performing injection attack method in various datasets, demonstrating its superiority and transferability. Peican Zhu, Zechen Pan, Keke Tang, Jinhuan Wang, Qi Xuan 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | A Novel Fuzzy Neural Network Architecture Search Framework for Defect Recognition With UncertaintiesabstractDefect recognition is an important task in intelligent manufacturing. Due to the subjectivity of human annotation, the collected defect data usually contains a lot of noise and unpredictable uncertainties, which have a great negative influence on defect recognition. It is a significant challenge to discover an effective defect recognition model with satisfactory uncertainty processing ability. A natural way is to automatically search for an efficient deep model, which can be realized by neural architecture search (NAS). To achieve this, we propose an efficient fuzzy NAS framework for defect recognition, where the searched architecture can effectively handle uncertain information from the given datasets. Specifically, we first design a fuzzy search space and the related encoding strategy for fuzzy NAS. Then, we propose a comparator-based evolutionary search approach, where an online end-to-end comparator is learned to directly determine the selection of candidate architectures from the evolutionary population. The comparator works in an end-to-end way and it transforms the complex ranking problem of evaluating architectures into a simple classification task, which overcomes the rank disorder issue suffered from traditional performance predictors. A series of experimental results demonstrate that the architecture with fewer #Params (1.22 M) search by fuzzy neural architecture search framework for defect recognition method achieves higher accuracy (92.26%) compared to the state-of-the-art results (i.e., DARTS-PV) on the ELPV dataset, as well as competitive results (accuracy = 76.4%, #Params = 1.04 M) on the CODEBRIM dataset. Experimental results show the effectiveness and efficiency of our proposed method in handling uncertain problems. Lianbo Ma 0004, Nan Li 0033, Peican Zhu, Keke Tang, Feng Wang 0048, Guo Yu 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | SelfGCN: Graph Convolution Network With Self-Attention for Skeleton-Based Action RecognitionabstractGraph Convolutional Networks (GCNs) are widely used for skeleton-based action recognition and achieved remarkable performance. Due to the locality of graph convolution, GCNs can only utilize short-range node dependencies but fail to model long-range node relationships. In addition, existing graph convolution based methods normally use a uniform skeleton topology for all frames, which limits the ability of feature learning. To address these issues, we present the Graph Convolution Network with Self-Attention (SelfGCN), which consists of a mixing features across self-attention and graph convolution (MFSG) module and a temporal-specific spatial self-attention (TSSA) module. The MFSG module models local and global relationships between joints by executing graph convolution and self-attention branches in parallel. Its bi-directional interactive learning strategy utilizes complementary clues in the channel dimensions and the spatial dimensions across both of these branches. The TSSA module uses self-attention to learn the spatial relationships between joints of each frame in a skeleton sequence. It also models the unique spatial features of the single frames. We conduct extensive experiments on three popular benchmark datasets, NTU RGB+D, NTU RGB+D120, and Northwestern-UCLA. The results of the experiment demonstrate that our method achieves or exceeds the record accuracies on all three benchmarks. Our project website is available at https://github.com/SunPengP/SelfGCN. Zhize Wu, Keke Tang, Tong Xu 0001, Le Zou, Xiaofeng Wang 0009, Fan Cheng 0001, Thomas Weise 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Multi-Resolution Wavelet Fractal Analysis and Subtask Training for Enhancing Few-Shot Noisy Brainwave RecognitionabstractThe integration of healthcare monitoring with Internet of Things (IoT) networks radically transforms the management and monitoring of human well-being. Portable and lightweight electroencephalography (EEG) systems with fewer electrodes have improved convenience and flexibility while retaining adequate accuracy. However, challenges emerge when dealing with real-time EEG data from IoT devices due to the presence of noisy samples, which impedes improvements in brainwave detection accuracy. Moreover, high inter-subject variability and substantial variability in EEG signals present difficulties for conventional data augmentation and subtask learning techniques, leading to poor generalizability. To address these issues, we present a novel framework for enhancing EEG-based recognition through multi-resolution data analysis, capturing features at different scales using wavelet fractals. The original data can be expanded many times after continuous wavelet transform (CWT) and recombination, alleviating insufficient training samples. In the transfer stage of deep learning (DL) models, we adopt a subtask learning approach to train the recognition model to generalize efficiently. This incorporates wavelets at various scales instead of exclusively considering average prediction performance across scales and paradigms. Through extensive experiments, we demonstrate that our proposed DL-based method excels at extracting features from small-scale and noisy EEG data. This significantly improves healthcare monitoring performance by mitigating the impact of noise introduced by the external environment. Denghui Zhang 0001, Muhammad Shafiq 0003, Keke Tang, Usman Naseem |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Rethinking Video Sentence Grounding From a Tracking Perspective With Memory Network and Masked AttentionabstractVideo sentence grounding (VSG) is the task of identifying the segment of an untrimmed video that semantically corresponds to a given natural language query. While many existing methods extract frame-grained features using pre-trained 2D or 3D convolution networks, often fail to capture subtle differences between ambiguous adjacent frames. Although some recent approaches incorporate object-grained features using Faster R-CNN to capture more fine-grained details, they are still primarily based on feature enhancement and lack spatio-temporal modeling to explore the semantics of the core persons/objects. To solve the problem of modeling the core target's behavior, in this paper, we propose a new perspective for addressing the VSG task by tracking pivotal objects and activities to learn more fine-grained spatio-temporal features. Specifically, we introduce the Video Sentence Tracker with Memory Network and Masked Attention (VSTMM), which comprises a cross-modal targets generator for producing multi-modal templates and search space, a memory-based tracker for dynamically tracking multi-modal targets using a memory network to record targets' behaviors, a masked attention localizer which learns local shared features between frames and eliminates interference from long-term dependencies, resulting in improved accuracy when localizing the moment. To evaluate the performance of our VSTMM, we conducted extensive experiments and comparisons with state-of-the-art methods on three challenging benchmarks, including Charades-STA, ActivityNet Captions, and TACoS. Without bells and whistles, our VSTMM achieves leading performance with a considerable real-time speed. Zeyu Xiong, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Jiahao Zhu 0003, Keke Tang, Pan Zhou 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Deep Manifold Attack on Point Clouds via Parameter Plane StretchingabstractAdversarial attack on point clouds plays a vital role in evaluating and improving the adversarial robustness of 3D deep learning models. Current attack methods are mainly applied by point perturbation in a non-manifold manner. In this paper, we formulate a novel manifold attack, which deforms the underlying 2-manifold surfaces via parameter plane stretching to generate adversarial point clouds. First, we represent the mapping between the parameter plane and underlying surface using generative-based networks. Second, the stretching is learned in the 2D parameter domain such that the generated 3D point cloud fools a pretrained classifier with minimal geometric distortion. Extensive experiments show that adversarial point clouds generated by manifold attack are smooth, undefendable and transferable, and outperform those samples generated by the state-of-the-art non-manifold ones. Keke Tang, Jianpeng Wu, Weilong Peng, Yawen Shi, Peng Song 0001, Zhaoquan Gu, Zhihong Tian 0001, Wenping Wang 0001 |
AAAI | 1 |
| 2023 | HEPT Attack: Heuristic Perpendicular Trial for Hard-label Attacks under Limited Query BudgetsabstractExploring adversarial attacks on deep neural networks (DNNs) is crucial for assessing and enhancing their adversarial robustness. Among various attack types, hard-label attacks that rely only on predicted labels offer a practical approach. This paper focuses on the challenging task of hard-label attacks within an extremely limited query budget, which is a significant achievement rarely accomplished by existing methods. To tackle this, we propose an attack framework that leverages geometric information from previous perturbation directions to form triangles and employs a heuristic perpendicular trial to effectively utilize the intermediate directions. Extensive experiments validate the effectiveness of our approach under strict query constraints and demonstrate its superiority to the state-of-the-art methods. Qi Li 0048, Keke Tang, Peican Zhu |
CIKM | 4 |
| 2023 | Matching Words for Out-of-distribution DetectionabstractDeep neural networks often exhibit the overconfidence issue when encountering out-of-distribution (OOD) samples. To address this, leveraging large-scale pre-trained models like CLIP has shown promise. While CLIP has the capability to encode a vast array of interconnected concepts, current OOD detection methods based on it primarily focus on ID categories and a limited set of OOD categories. In this paper, we propose a novel approach that harnesses the power of WordNet to fully exploit the rich knowledge encapsulated within CLIP, resulting in enhanced OOD detection performance. Our methodology involves constructing a word tree that includes both in-distribution (ID) words and a large set of semantically similar OOD words selected from WordNet. By matching a test image with the concepts of the words in the word tree using CLIP, we estimate the probability of the image being classified as either ID or OOD. Furthermore, we introduce a conditional random field model to effectively handle both the parent-child and the sibling-sibling conflicts in the concept matching results. Extensive experiments under various ID/OOD settings demonstrate the effectiveness of our approach and its superiority over state-of-the-art methods. Keke Tang, Xujian Cai, Weilong Peng, Daizong Liu, Peican Zhu, Pan Zhou 0001, Zhihong Tian 0001, Wenping Wang 0001 |
ICDM | 1 |
| 2023 | OOD Attack: Generating Overconfident out-of-Distribution Examples to Fool Deep Neural ClassifiersabstractDeep neural networks (DNNs) are dominating various computer vision solutions. However, DNN classifiers suffer from the out-of-distribution (OOD) overconfidence issue, i.e., making overconfident predictions on OOD samples. In this paper, we consider a new OOD attack task, i.e., generating OOD examples that fool DNN classifiers to trap into this issue. Specifically, we first generate seed examples by sampling from common OOD distributions, and then lift the prediction to be overconfident. Extensive experiments with different seeds and confidence-lifting solutions under white-and black-box settings validate the feasibility of OOD attack. Besides, we demonstrate its usefulness in evaluating OOD detection and alleviating the OOD overconfidence issue. Keke Tang, Xujian Cai, Weilong Peng, Shudong Li, Wenping Wang 0001 |
ICIP | 1 |
| 2023 | DBA: An Efficient Approach to Boost Transfer-Based Adversarial Attack Performance Through Information Deletion
Zepeng Fan, Peican Zhu, Chao Gao 0001, Jinbang Hong, Keke Tang |
KSEM (2) | 5 |
| 2023 | Enhancing Adversarial Robustness via Anomaly-aware Adversarial Training
Keke Tang, Tianrui Lou, Yawen Shi, Peican Zhu, Zhaoquan Gu |
KSEM (1) | 1 |
| 2023 | Are Deep Point Cloud Classifiers Suffer From Out-of-distribution Overconfidence Issue?abstract3D point cloud perception using deep neural networks (DNNs) has been a trend for various application scenarios. However, the black-box nature of DNNs will bring many hidden risks as in the 2D image field. In this paper, we present a preliminary evaluation on the out-of-distribution (OOD) overconfidence issue of deep point cloud classifiers, which has been proven to exist in deep 2D image classifiers, i.e., OOD inputs will lead to overconfident predictions on predefined categories. We also investigate whether a simple thresholding baseline and two modern OOD detection solutions can handle the issue by detecting OOD samples. Extensive experiments with four representative deep point cloud classifiers train/evaluate on different in/out-of-distribution point clouds validate the severity and knottiness of the OOD overconfidence issue. Our investigation will provide the groundwork for future studies on handling the OOD overconfidence issue of DNN classifiers for 3D point clouds. Keke Tang, Yawen Shi, Weilong Peng, Peican Zhu |
SMC | 2 |
| 2023 | Rethinking Perturbation Directions for Imperceptible Adversarial Attacks on Point CloudsabstractAdversarial attacks have been successfully extended to the field of point clouds. Besides applying the common perturbation guided by the gradient, adversarial attacks on point clouds can be conducted by applying directional perturbations, e.g., along normal and along the tangent plane. In this article, we first investigate whether adversarial attacks with these two orthogonal directional perturbations are more imperceptible than that with the gradient-aware perturbation. Second, we investigate the deeper difference between adversarial attacks with these two directional perturbations, and whether they are applicable to the same scenarios. Third, based on the verification results that the above two directional perturbations have different sensitiveness to curvature, we devise a novel normal-tangent attack (NTA) framework with a hybrid directional perturbation scheme that adaptively chooses the direction according to the curvature of the local shape around the point. Extensive experiments on two publicly available data sets, e.g., ModelNet40 and ShapeNet Part, with classifiers in three representative networks, e.g., PointNet++, DGCNN, PointConv, validate the effectiveness of NTA, and the superiority to the state-of-the-art methods. Keke Tang, Yawen Shi, Tianrui Lou, Weilong Peng, Peican Zhu, Zhaoquan Gu, Zhihong Tian 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Unsupervised feature selection through combining graph learning and ℓ2,0-norm constraint
Peican Zhu, Keke Tang, Yang Liu 0144, Yin-Ping Zhao, Zhen Wang 0004 |
Inf. Sci. | 3 |
| 2023 | RepPVConv: attentively fusing reparameterized voxel features for efficient 3D point cloud perception
Keke Tang, Weilong Peng, Yanling Zhang, Meie Fang, Zheng Wang 0002, Peng Song 0001 |
Vis. Comput. | 1 |
| 2022 | GM-Attack: Improving the Transferability of Adversarial Attacks
Jinbang Hong, Keke Tang, Chao Gao 0001, Songxin Wang, Sensen Guo, Peican Zhu |
KSEM (3) | 2 |
| 2021 | CODEs: Chamfer Out-of-Distribution Examples against Overconfidence IssueabstractOverconfident predictions on out-of-distribution (OOD) samples is a thorny issue for deep neural networks. The key to resolve the OOD overconfidence issue inherently is to build a subset of OOD samples and then suppress predictions on them. This paper proposes the Chamfer OOD examples (CODEs), whose distribution is close to that of in-distribution samples, and thus could be utilized to alleviate the OOD overconfidence issue effectively by suppressing predictions on them. To obtain CODEs, we first generate seed OOD examples via slicing&splicing operations on in-distribution samples from different categories, and then feed them to the Chamfer generative adversarial network for distribution transformation, without accessing to any extra data. Training with suppressing predictions on CODEs is validated to alleviate the OOD overconfidence issue largely without hurting classification accuracy, and outperform the state-of-the-art methods. Besides, we demonstrate CODEs are useful for improving OOD detection and classification. Keke Tang, Dingruibo Miao, Weilong Peng, Jianpeng Wu, Yawen Shi, Zhaoquan Gu, Zhihong Tian 0001, Wenping Wang 0001 |
ICCV | 1 |
| 2019 | Computational Design of Steady 3D Dissection PuzzlesabstractAbstract Dissection puzzles require assembling a common set of pieces into multiple distinct forms. Existing works focus on creating 2D dissection puzzles that form primitive or naturalistic shapes. Unlike 2D dissection puzzles that could be supported on a tabletop surface, 3D dissection puzzles are preferable to be steady by themselves for each assembly form. In this work, we aim at computationally designing steady 3D dissection puzzles. We address this challenging problem with three key contributions. First, we take two voxelized shapes as inputs and dissect them into a common set of puzzle pieces, during which we allow slightly modifying the input shapes, preferably on their internal volume, to preserve the external appearance. Second, we formulate a formal model of generalized interlocking for connecting pieces into a steady assembly using both their geometric arrangements and friction. Third, we modify the geometry of each dissected puzzle piece based on the formal model such that each assembly form is steady accordingly. We demonstrate the effectiveness of our approach on a wide variety of shapes, compare it with the state‐of‐the‐art on 2D and 3D examples, and fabricate some of our designed puzzles to validate their steadiness. Keke Tang, Peng Song 0001, Bailin Deng, Chi-Wing Fu, Ligang Liu 0001 |
Comput. Graph. Forum | 1 |
| 2016 | Signature of Geometric Centroids for 3D Local Shape Description and Partial Shape Matching
Keke Tang, Peng Song 0001 |
ACCV (5) | 1 |
| 2015 | KeJia-LC: A Low-Cost Mobile Robot Platform - Champion of Demo Challenge on Benchmarking Service Robots at RoboCup 2015abstractIn this paper, we present the system design and the key techniques of our mobile robot platform called KeJia-LC , who won the first place in the demo challenge on Benchmarkinng Service Robots in RoboCup 2015. Given the fact that KeJia-LC is a low-cost version of our KeJia robot without shoulder and arm, several new technical demands comparing to RoboCup@Home are highlighted for better understanding of our system. With the elaborate design of hardware and the reasonable selection of sensors, our robot platform has the features of low cost, wide generality and good extensibility. Moreover, we integrate several functional softwares (such as 2D&3D mapping, localization and navigation) following the competition rules, which are critical to the performance of our robot. The effectiveness and robustness of our robot system has been proven in the competition. Feng Wu 0001, Ningyang Wang, Keke Tang |
RoboCup | 4 |
| 2015 | Synthetical Benchmarking of Service Robots: A First Effort on Domestic Mobile PlatformsabstractMost of existing benchmarking tools for service robots are basically qualitative, in which a robot’s performance on a task is evaluated based on completion/incompletion of actions contained in the task. In the effort reported in this paper, we tried to implement a synthetical benchmarking system on domestic mobile platforms. Synthetical benchmarking consists of both qualitative and quantitative aspects, such as task completion, accuracy of task completions and efficiency of task completions, about performance of a robot. The system includes a set of algorithms for collecting, recording and analyzing measurement data from a MoCap system. It was used as the evaluator in a competition called the BSR challenge, in which 10 teams participated, at RoboCup 2015. The paper presents our motivations behind synthetical benchmarking, the design considerations on the synthetical benchmarking system, the realization of the competition as a comparative study on performance evaluation of domestic mobile platforms, and an analysis of the teams’ performance. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Keke Tang, Feng Wu 0001, Andras Gabor Kupcsik, Luca Iocchi, David Hsu |
RoboCup | 3 |
| 2014 | The Intelligent Techniques in Robot KeJia - The Champion of RoboCup@Home 2014
Dongcai Lu, Keke Tang, Ningyang Wang |
RoboCup | 4 |