EDBT 2026 Demo / reviewers in the wild / expert
Wenrui Li 0001
dblp:38/7677-1
· DBLP profile ↗
30ranked-venue papers
16as first author
29since 2021 · last 2026
0000-0002-2393-9016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 15 first-author · 24 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D RetrievalabstractWith the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution (RISD). HRC compresses abstract-to-specific and whole-to-part hierarchies in Euclidean embeddings, while RISD averages noisy fragments, obscuring critical semantic cues and diminishing the model’s ability to distinguish hard negatives. To address these challenges, we introduce the Hyperbolic Hierarchical Alignment Reasoning Network (H2ARN) for text-3D retrieval. H2ARN embeds both text and 3D data in a Lorentz-model hyperbolic space, where exponential volume growth inherently preserves hierarchical distances. A hierarchical ordering loss constructs a shrinking entailment cone around each text vector, ensuring that the matched 3D instance falls within the cone, while an instance-level contrastive loss jointly enforces separation from non-matching samples. To tackle RISD, we propose a contribution-aware hyperbolic aggregation module that leverages Lorentzian distance to assess the relevance of each local feature and applies contribution-weighted aggregation guided by hyperbolic geometry, enhancing discriminative regions while suppressing redundancy without additional supervision. We also release the expanded T3DR-HIT v2 benchmark, which contains 8,935 text-to-3D pairs, 2.6 times the original size, covering both fine-grained cultural artefacts and complex indoor scenes. Wenrui Li 0001, Yidan Lu, Yeyu Chai, Rui Zhao 0010, Hengyu Man, Xiaopeng Fan 0001 |
AAAI | 1 |
| 2026 | MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image CompressionabstractRecent advances in extreme image compression have revealed that mapping pixel data into highly compact latent representations can significantly improve coding efficiency. However, most existing methods compress images into 2-D latent spaces via convolutional neural networks (CNNs) or Swin Transformers, which tend to retain substantial spatial redundancy, thereby limiting overall compression performance. In this paper, we propose a novel Mixed RWKV-Transformer (MRT) architecture that encodes images into more compact 1-D latent representations by synergistically integrating the complementary strengths of linear-attention-based RWKV and self-attention-based Transformer models. Specifically, MRT partitions each image into fixed-size windows, utilizing RWKV modules to capture global dependencies across windows and Transformer blocks to model local redundancies within each window. The hierarchical attention mechanism enables more efficient and compact representation learning in the 1-D domain. To further enhance compression efficiency, we introduce a dedicated RWKV Compression Model (RCM) tailored to the structure characteristics of the intermediate 1-D latent features in MRT. Extensive experiments on standard image compression benchmarks validate the effectiveness of our approach. The proposed MRT framework consistently achieves superior reconstruction quality at bitrates below 0.02 bits per pixel (bpp). Quantitative results based on the DISTS metric show that MRT significantly outperforms the state-of-the-art 2-D architecture GLC, achieving bitrate savings of 43.75%, 30.59% on the Kodak and CLIC2020 test datasets, respectively. Hengyu Man, Wenrui Li 0001, Debin Zhao |
AAAI | 4 |
| 2026 | T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low BitratesabstractRecent advances in video generation techniques have given rise to an emerging paradigm of generative video coding for Ultra-Low Bitrate (ULB) scenarios by leveraging powerful generative priors. However, most existing methods are limited by domain specificity (e.g., facial or human videos) or excessive dependence on high-level text guidance, which tend to inadequately capture fine-grained motion details, leading to unrealistic or incoherent reconstructions. To address these challenges, we propose Trajectory-Guided Generative Video Coding (dubbed T-GVC), a novel framework that bridges low-level motion tracking with high-level semantic understanding. T-GVC features a semantic-aware sparse motion sampling pipeline that extracts pixel-wise motion as sparse trajectory points based on their semantic importance, significantly reducing the bitrate while preserving critical temporal semantic information. In addition, by integrating trajectory-aligned loss constraints into diffusion processes, we introduce a training-free guidance mechanism in latent space to ensure physically plausible motion patterns without sacrificing the inherent capabilities of generative models. Experimental results demonstrate that T-GVC outperforms both traditional and neural video codecs under ULB conditions. Furthermore, additional experiments confirm that our framework achieves more precise motion control than existing text-guided methods, paving the way for a novel direction of generative video coding guided by geometric motion modeling. Zhitao Wang, Hengyu Man, Wenrui Li 0001, Xiaopeng Fan 0001, Debin Zhao |
AAAI | 3 |
| 2026 | Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero-shot learning
Siteng Ma, Haocheng Tang, Jisheng Chu, Wenrui Li 0001 |
Neurocomputing | 6 |
| 2026 | Language-Guided Graph Representation Learning for Video SummarizationabstractWith the rapid growth of video content on social media, video summarization has become a crucial task in multimedia processing. However, existing methods face challenges in capturing global dependencies in video content and accommodating multimodal user customization. Moreover, temporal proximity between video frames does not always correspond to semantic proximity. To tackle these challenges, we propose a novel Language-guided Graph Representation Learning Network (LGRLN) for video summarization. Specifically, we introduce a video graph generator that converts video frames into a structured graph to preserve temporal order and contextual dependencies. By constructing forward, backward and undirected graphs, the video graph generator effectively preserves the sequentiality and contextual relationships of video content. We designed an intra-graph relational reasoning module with a dual-threshold graph convolution mechanism, which distinguishes semantically relevant frames from irrelevant ones between nodes. Additionally, our proposed language-guided cross-modal embedding module generates video summaries with specific textual descriptions. We model the summary generation output as a mixture of Bernoulli distribution and solve it with the EM algorithm. Experimental results show that our method outperforms existing approaches across multiple benchmarks. Moreover, we proposed LGRLN reduces inference time and model parameters by 87.8% and 91.7%, respectively. Wenrui Li 0001, Wei Han 0002, Hengyu Man, Wangmeng Zuo, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Hierarchical Neural Skill-Based Meta-Reinforcement Learning for Efficient Adaptability in Robotic Manipulation TasksabstractRobotic manipulation tasks frequently share foundational structures. Meta-reinforcement learning aims to develop generalizable policies that leverage these shared structures. However, existing methods usually struggle to efficiently encode this task-specific knowledge into their policy networks: hierarchical policies depend on task-agnostic action-level skills or intrinsic rewards, while context-based paradigms suffer from Markov Decision Process ambiguity. To address these challenges, we propose a hierarchical neural skill-based meta-reinforcement learning framework. This framework includes neural skill generation, neural skill decoding, and policy network construction. A Transformer-based neural skill generation unit sequentially generates hierarchical neural skills conditioned on a task description. The decoding mechanism translates these neural skills into network parameters for a policy. Using these decoded parameters, the policy network is constructed layer by layer to process the state information. Unlike previous work, our method treats skills as abstractions of layer-wise network parameters, allowing task-specific knowledge embedded in neural skills to directly configure the policy network. Experimental results demonstrate that our method has enhanced flexibility and efficiency. Hao Wang 0212, Wenrui Li 0001, Penghong Wang, Xianqi Zhang, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal RetrievalabstractCross-modal retrieval is essential for exploring semantic correlations between multimodal data. However, existing approaches face challenges in resolving semantic ambiguity and transferring knowledge with sparse sample generalization. To address these challenges, we propose a new Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network (SKPMN). Specifically, the Semantic Decoupling and Distinction (SDD) module decomposes complex word-region relationships into relevance-driven representations. The Deep Probability Mapping (DPM) module introduces a paradigm shift by mapping multimodal features into probabilistic distributions, capturing the semantic similarities and the potential uncertainties that define sparse or ambiguous relationships. By combining the Attention Probabilistic Mapping (APM) module, the model can effectively transfer knowledge across similar samples while emphasizing critical distinctions, significantly enhancing generalization to sparse and ambiguous samples. Finally, the multi-grained alignment strategy establishes a novel integration of fine-grained patch-to-word alignment and coarse-grained global alignment. Experimental results show that SKPMN achieves superior retrieval accuracy across major benchmark datasets. Furthermore, we implement a channel resource allocation technique that allocates more transmission resources to semantically significant information. In resource-constrained environments, our approach leverages Joint Source-Channel Coding (JSCC) to enhance the efficiency of visual feature transmission. Wenrui Li 0001, Yeyu Chai, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | DV-Hop Localization Based on Probability Distance Estimation and Expected Hop Distance CorrectionabstractDistance estimation and theoretical derivation in 3D space form the foundation basis for improving localization performance in wireless sensor networks (WSNs). Localization is a pivotal challenge in wireless sensor network (WSN) applications. To address this issue, we propose a probability-based distance estimation (PDE) model and a distance correction strategy based on expected hops (DCSEH). First, the PDE model is constructed from the multi-hop probability distribution of nodes, from which the upper bound and average distance for anchor nodes to detect target nodes under different hop counts are derived. Second, the DCSEH strategy effectively mitigates transmission-path detours in wireless node communication. Finally, the constructed loss function is embedded into a multi-objective genetic algorithm to predict the position of each unknown node in three-dimensional space. Extensive experiments demonstrate that the proposed method achieves state-of-the-art 3D localization performance on both random and multimodal datasets. Penghong Wang, Hao Wang 0212, Wenrui Li 0001, Hengyu Man, Xin Yue, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Digging into Intrinsic Contextual Information for High-fidelity 3D Point Cloud CompletionabstractThe common occurrence of occlusion-induced incompleteness in point clouds has made point cloud completion (PCC) a highly-concerned task in the field of geometric processing. Existing PCC methods typically produce complete point clouds from partial point clouds in a coarse-to-fine paradigm, with the coarse stage generating entire shapes and the fine stage improving texture details. Though diffusion models have demonstrated effectiveness in the coarse stage, the fine stage still faces challenges in producing high-fidelity results due to the ill-posed nature of PCC. The intrinsic contextual information for texture details in partial point clouds is the key to solving the challenge. In this paper, we propose a high-fidelity PCC method that digs into both short and long-range contextual information from the partial point cloud in the fine stage. Specifically, after generating the coarse point cloud via a diffusion-based coarse generator, a mixed sampling module introduces short-range contextual information from partial point clouds into the fine stage. A surface freezing module safeguards points from noise-free partial point clouds against disruption. As for the long-range contextual information, we design a similarity modeling module to derive similarity with rigid transformation invariance between points, conducting effective matching of geometric manifold features globally. In this way, the high-quality components present in the partial point cloud serve as valuable references to refine the coarse point cloud with high fidelity. Extensive experiments have demonstrated the superiority of the proposed method over SOTA competitors. Jisheng Chu, Wenrui Li 0001, Kanglin Ning, Yidan Lu, Xiaopeng Fan 0001 |
AAAI | 2 |
| 2025 | Riemann-based Multi-scale Attention Reasoning Network for Text-3D RetrievalabstractDue to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemann-based Multi-scale Attention Reasoning Network (RMARN) for text-3D retrieval. Specifically, the extracted text and point cloud features are refined by their respective Adaptive Feature Refiner (AFR). Furthermore, we introduce the innovative Riemann Local Similarity (RLS) module and the Global Pooling Similarity (GPS) module. However, as 3D point cloud data and text data often possess complex geometric structures in high-dimensional space, the proposed RLS employs a novel Riemann Attention Mechanism to reflect the intrinsic geometric relationships of the data. Without explicitly defining the manifold, RMARN learns the manifold parameters to better represent the distances between text-point cloud samples. To address the challenges of lacking paired text-3D data, we have created the large-scale Text-3D Retrieval dataset T3DR-HIT, which comprises over 3,380 pairs of text and point cloud data. T3DR-HIT contains coarse-grained indoor 3D scenes and fine-grained Chinese artifact scenes, consisting of 1,380 and over 2,000 text-3D pairs, respectively. Experiments on our custom datasets demonstrate the superior performance of the proposed method. Wenrui Li 0001, Wei Han 0002, Yandu Chen, Yeyu Chai, Yidan Lu, Xiaopeng Fan 0001 |
AAAI | 1 |
| 2025 | Hyperbolic-Constraint Point Cloud Reconstruction from Single RGB-D ImagesabstractReconstructing desired objects and scenes has long been a primary goal in 3D computer vision. Single-view point cloud reconstruction has become a popular technique due to its low cost and accurate results. However, single-view reconstruction methods often rely on expensive CAD models and complex geometric priors. Effectively utilizing prior knowledge about the data remains a challenge. In this paper, we introduce hyperbolic space to 3D point cloud reconstruction, enabling the model to represent and understand complex hierarchical structures in point clouds with low distortion. We build upon previous methods by proposing a hyperbolic Chamfer distance and a regularized triplet loss to enhance the relationship between partial and complete point clouds. Additionally, we design adaptive boundary conditions to improve the model's understanding and reconstruction of 3D structures. Our model outperforms most existing models, and ablation studies demonstrate the significance of our model and its components. Experimental results show that our method significantly improves feature extraction capabilities. Our model achieves outstanding performance in 3D reconstruction tasks. Wenrui Li 0001, Wei Han 0002, Hengyu Man, Xiaopeng Fan 0001 |
AAAI | 1 |
| 2025 | Text-Guided Editable 3D City Scene GenerationabstractThe automated generation of 3D city scenes has attracted considerable attention due to its broad applications in areas such as virtual reality, urban planning, and digital media. Traditional approaches for constructing 3D city environments typically depend on labor-intensive manual modeling or the use of complex, non-editable training models. To overcome these limitations, we propose an innovative framework that generates fully editable 3D city scenes directly from natural language descriptions. Our framework utilizes a structured data extraction process to decouple model and layout features from textual descriptions, facilitating the creation of 2D layouts that guide the generation of 3D terrains. Furthermore, we introduce a constrained 3D terrain generation method that ensures consistency with the semantic content and spatial relationships delineated in the input text. By integrating 2D layouts, 3D terrains, and procedural modeling techniques, our framework creates a fully editable 3D environment, empowering users to efficiently customize and modify scene components. Experimental results indicate that our method offers substantial enhancements in flexibility, realism, and user-friendliness, positioning it as a promising approach to democratize 3D city modeling. Yuchuan Feng, Jihang Jiang, Wenrui Li 0001, Ruotong Li, Xiaopeng Fan 0001 |
ICASSP | 4 |
| 2025 | Discrepancy-Aware Attention Network for Enhanced Audio-Visual Generalized Zero-Shot LearningabstractAudio-visual Generalized Zero-Shot Learning ((G)ZSL) has attracted significant attention for its ability to identify unseen classes in general video classification tasks. However, modality imbalance in (G)ZSL leads to over-reliance on the optimal modality, reducing discriminative capabilities for unseen classes. Though recent studies have attempted to address this issue, two challenges still remain unsolved: (a) Quality discrepancies, where modalities offer differing quantities and qualities of information for the same concept. (b) Content discrepancies, where the contributions of different samples within the same modality exhibit significant differences. To address these challenges, we propose a Discrepancy-Aware Attention Network (DAAN) for Enhanced Audio-Visual (G)ZSL. Our approach introduces a Redundant-Noise Mitigation Attention (RNMA) unit to minimize content discrepancies by mitigating redundant information in modalities and a Contrastive Sample Gradient Modulation (CSGM) mechanism to adjust gradient magnitudes and balance quality discrepancies. We quantify modality contributions by integrating optimization and convergence rate for more precise gradient modulation in CSGM. Experiments demonstrate DAAN achieves state-of-the-art performance on benchmark datasets, with ablation studies validating the effectiveness of individual modules. Code is available at https://github.com/xiaoxinning/DAAN-GZSL. Runlin Yu, Yipu Gong, Wenrui Li 0001, Aiwen Sun, Mengren Zheng |
ACM Multimedia | 3 |
| 2025 | Multi-modal spiking tensor regression network for audio-visual zero-shot learning
Wenrui Li 0001, Jinxiu Hou, Guanghui Cheng |
Neurocomputing | 2 |
| 2025 | Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot LearningabstractAudio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scene biases and inadequate motion detail. This paper proposes a novel dual-stream Multi-Timescale Motion-Decoupled Spiking Transformer (MDST++), which decouples contextual semantic information and sparse dynamic motion information. The recurrent joint learning unit is proposed to extract contextual semantic information and capture joint knowledge across various modalities to understand the environment of actions. By converting RGB images to events, our method captures motion information more accurately and mitigates background scene biases. Moreover, we introduce a discrepancy analysis block to model audio motion information. To enhance the robustness of SNNs in extracting temporal and motion cues, we dynamically adjust the threshold of Leaky Integrate-and-Fire neurons based on global motion and contextual semantic information. Our experiments validate the effectiveness of MDST++, demonstrating their consistent superiority over state-of-the-art methods on mainstream benchmarks. Additionally, incorporating motion and multi-timescale information significantly improves HM and ZSL accuracy by 26.2% and 39.9%. Wenrui Li 0001, Penghong Wang, Wangmeng Zuo, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Spiking Variational Graph Representation Inference for Video SummarizationabstractWith the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic coherence of video content. Additionally, these methods are also influenced by noise during multi-channel feature fusion. We propose a Spiking Variational Graph (SpiVG) Network, which enhances information density and reduces computational complexity. First, we design a keyframe extractor based on Spiking Neural Networks (SNN), leveraging the event-driven computation mechanism of SNNs to learn keyframe features autonomously. To enable fine-grained and adaptable reasoning across video frames, we introduce a Dynamic Aggregation Graph Reasoner, which decouples contextual object consistency from semantic perspective coherence. We present a Variational Inference Reconstruction Module to address uncertainty and noise arising during multi-channel feature fusion. In this module, we employ Evidence Lower Bound Optimization (ELBO) to capture the latent structure of multi-channel feature distributions, using posterior distribution regularization to reduce overfitting. Experimental results show that SpiVG surpasses existing methods across multiple datasets such as SumMe, TVSum, VideoXum, and QFVS. Our codes and pre-trained models are available at https://github.com/liwrui/SpiVG. Wenrui Li 0001, Wei Han 0002, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multi-Scale Spiking Pyramid Wireless Communication Framework for Food RecognitionabstractFood recognition applications in human health have recently garnered significant attention in the field of computer vision. With the advancement of mobile devices, robust food recognition in wireless communication has become a practical and challenging application scenario. We propose a novel Multi-scale Spiking Pyramid Transmission Network (MSPTN) to tackle this challenge. The MSPTN learns diverse and complementary local and global feature maps simultaneously, generating a comprehensive description of food images that capture the correlations of feed-specific features. The feature sender uses a three-layer Spiking Neural Network (SNN). The proposed sender compresses features into sparse and discrete spike trains, significantly reducing the required transmission bandwidth and improving channel utilization and energy efficiency. Our model introduces the Compressed Factorized Bilinear block (CFB), which employs a low-rank feature approximation to reduce computational complexity and feature transmission volume while preserving the discriminate features. The enhancement reasoning module is proposed to enhance the received features by projecting them into a higher-dimensional space and utilizing the self-attention mechanism and sum pooling to compress them back to the original dimension. We conduct extensive experiments on the ETH Food-101 and Food2k datasets. Our results reveal that the MSPTN demonstrates state-of-the-art recognition performance, even with binary spike trains. Meanwhile, the MSPTN also exhibits remarkable robustness in wireless communication scenarios. With the combination of CFB, SNN, and EFB, our model achieves significant efficiency gains, including a nearly nine-fold decrease in feature transmission volume and a three-fold improvement in runtime & computational memory speed. Wenrui Li 0001, Jiahui Li 0001, Mengyao Ma, Xiaopeng Hong, Xiaopeng Fan 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Smile: Spiking Multi-Modal Interactive Label-Guided Enhancement Network for Emotion RecognitionabstractMulti-modal multi-label emotion recognition has gained significant attention in the field of affective computing, enabling various signals to distinguish complex emotions accurately. However, previous studies primarily focus on capturing invariant representations, neglecting the importance of incorporating the fluctuation of temporal information which affects the model robustness. In this paper, we propose a novel Spiking Multi-modal Interactive Label-guided Enhancement network (SMILE). It introduces the spiking neural network with dynamic thresholds, allowing flexible processing of temporal information to enhance the model robustness. Furthermore, it employs the scale spiking fusion to enrich semantic information. In addition to modality-specific refinement, SMILE integrates the modality-interactive exploration and label-modality matching modules to capture multimodal interaction and label-modality dependence. Experimental results on benchmark datasets CMU-MOSEI and NEMu demonstrate the superiority of SMILE over state-ofthe-art models. Notably, SMILE achieves a significant 28.5% improvement in accuracy compared to the benchmark method when evaluated on NEMu dataset. Wenrui Li 0001, Yuxin Ge, Chong-Jun Wang |
ICME | 2 |
| 2024 | Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training ModelsabstractRecent studies on AI security have highlighted the vulnerability of Vision-Language Pre-training (VLP) models to subtle yet intentionally designed perturbations in images and texts. Investigating multimodal systems' robustness via adversarial attacks is crucial in this field. Most multimodal attacks are sample-specific, generating a unique perturbation for each sample to construct adversarial samples. To the best of our knowledge, it is the first work through multimodal decision boundaries to explore the creation of a universal, sample-agnostic perturbation that applies to any image. Initially, we explore strategies to move sample points beyond the decision boundaries of linear classifiers, refining the algorithm to ensure successful attacks under the top k accuracy metric. Based on this foundation, in visual-language tasks, we treat visual and textual modalities as reciprocal sample points and decision hyperplanes, guiding image embeddings to traverse text-constructed decision boundaries, and vice versa. This iterative process consistently refines a universal perturbation, ultimately identifying a singular direction within the input space which is exploitable to impair the retrieval performance of VLP models. The proposed algorithms support the creation of global perturbations or adversarial patches. Comprehensive experiments validate the effectiveness of our method, showcasing its data, task, and model transferability across various VLP models and datasets. Code: https://github.com/LibertazZ/MUAP Haonan Zheng 0001, Wen Jiang 0002, Xinyang Deng, Wenrui Li 0001 |
ACM Multimedia | 4 |
| 2024 | A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsabstractWith Vision-Language Pre-training (VLP) models demonstrating powerful multimodal interaction capabilities, the application scenarios of neural networks are no longer confined to unimodal domains but have expanded to more complex multimodal V+L downstream tasks. The security vulnerabilities of unimodal models have been extensively examined, whereas those of VLP models remain challenging. We note that in CV models, the understanding of images comes from annotated information, while VLP models are designed to learn image representations directly from raw text. Motivated by this discrepancy, we developed the Feature Guidance Attack (FGA), a novel method that uses text representations to direct the perturbation of clean images, resulting in the generation of adversarial images. FGA is orthogonal to many advanced attack strategies in the unimodal domain, facilitating the direct application of rich research findings from the unimodal to the multimodal scenario. By appropriately introducing text attack into FGA, we construct Feature Guidance with Text Attack (FGA-T). Through the interaction of attacking two modalities, FGA-T achieves superior attack effects against VLP models. Moreover, incorporating data augmentation and momentum mechanisms significantly improves the black-box transferability of FGA-T. Our method demonstrates stable and effective attack capabilities across various datasets, downstream tasks, and both black-box and white-box settings, offering a unified baseline for exploring the robustness of VLP models. Haonan Zheng 0001, Xinyang Deng, Wen Jiang 0002, Wenrui Li 0001 |
ACM Multimedia | 4 |
| 2024 | DV-Hop Localization Based on Distance Estimation Using Multinode and Hop Loss in IoTabstractSensor location awareness is a critical issue in internet of things applications. For more accurate location estimation, the two issues should be considered extensively: 1) how to sufficiently utilize the connection information between multiple nodes and 2) how to select a suitable solution from multiple solutions obtained by the Euclidean distance loss. In this paper, a DV-Hop localization based on the distance estimation using multi-node (DEMN) and the hop loss in WSNs is proposed to address the two issues. In DEMN, when multiple anchor nodes can detect an unknown node, the distance expectation between the unknown node and an anchor node is calculated using the cross domain information and is considered as the expected distance between them, which narrows the search space. When minimizing the traditional Euclidean distance loss, multiple solutions may exist. To select a suitable solution, the hop loss is proposed, which minimizes the difference between the real and its predicted hops. Finally, the Euclidean distance loss calculated by the DEMN and the hop loss are embedded into the multi-objective optimization algorithm. The experimental results show that the proposed method gains 86.11% location accuracy in the randomly distributed network, which is 6.05% better than the DEM-DV-Hop, while DEMN and the hop loss can contribute 2.46% and 3.41%, respectively. Penghong Wang, Wenrui Li 0001, Xiaopeng Fan 0001, Debin Zhao |
IEEE Internet Things J. | 3 |
| 2024 | Multi-Layer Probabilistic Association Reasoning Network for Image-Text RetrievalabstractWith the advancement of deep learning, the task of image-text retrieval has received widespread attention for addressing the semantic heterogeneity in multimodal data. However, many existing methods ignore the uncertainty present in manually annotated datasets. It is crucial for models to learn the potential corresponding relationships between regions in images and words in sentences. To tackle these challenges, we introduce the Multi-layer Probabilistic Association Reasoning Network (MPARN). In MPARN, the region-word association reasoning module is developed to treat each visual and textual fragment as unique probability distributions. This allows our model to imagine and capture the intricate one-to-many and many-to-many relationships between visual and textual objects. To effectively integrate the association distributions between visual and textual modalities, we propose the cross-modal association probability composer. This composer not only combines these distributions effectively but also preserves the intrinsic hierarchical structure of the elements involved. Furthermore, we introduce the semantic relationship reasoning module, which is designed to analyze the contextual semantic information within each modality. The multi-layer adaptive aggregate composer is employed to progressively explore semantic correlations within each modality and to dynamically synthesize outputs based on their relevance. Our extensive experiments on the Flickr30K and MSCOCO datasets demonstrate the MPARN’s state-of-the-art retrieval performance when compared to other baselines. The qualitative results further validate the effectiveness of the probabilistic association distributions. Wenrui Li 0001, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot LearningabstractThe spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling SNNs (binary spike sequences) with transformers (float-point sequences) to jointly explore the temporal-semantic information still facing challenges. In this paper, we introduce a novel Spiking Tucker Fusion Transformer (STFT) for audio-visual zero-shot learning (ZSL). The STFT leverage the temporal and semantic information from different time steps to generate robust representations. The time-step factor (TSF) is introduced to dynamically synthesis the subsequent inference information. To guide the formation of input membrane potentials and reduce the spike noise, we propose a global-local pooling (GLP) which combines the max and average pooling operations. Furthermore, the thresholds of the spiking neurons are dynamically adjusted based on semantic and temporal cues. Integrating the temporal and semantic information extracted by SNNs and Transformers are difficult due to the increased number of parameters in a straightforward bilinear model. To address this, we introduce a temporal-semantic Tucker fusion module, which achieves multi-scale fusion of SNN and Transformer outputs while maintaining full second-order interactions. Our experimental results demonstrate the effectiveness of the proposed approach in achieving state-of-the-art performance in three benchmark datasets. The harmonic mean (HM) improvement of VGGSound, UCF101 and ActivityNet are around 15.4%, 3.9%, and 14.9%, respectively. Wenrui Li 0001, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Modality-Fusion Spiking Transformer Network for Audio-Visual Zero-Shot LearningabstractAudio-visual zero-shot learning (ZSL), which learns to classify video data from the classes not being observed during training, is challenging. In audio-visual ZSL, both semantic and temporal information from different modalities is relevant to each other. However, effectively extracting and fusing information from audio and visual remains an open challenge. In this work, we propose an Audio-Visual Modality-fusion Spiking Transformer network (AVMST) for audio-visual ZSL. To be more specific, AVMST provides a spiking neural network (SNN) module for extracting conspicuous temporal information of each modality, a cross-attention block to effectively fuse the temporal and semantic information, and a transformer reasoning module to further explore the interrelationships of fusion features. To provide robust temporal features, the spiking threshold of the SNN module is adjusted dynamically based on the semantic cues of different modalities. The generated feature map is in accordance with the zero-shot learning property thanks to our proposed spiking transformer’s ability to combine the robustness of SNN feature extraction and the precision of transformer feature inference. Extensive experiments on three benchmark audio-visual datasets (i.e., VGGSound, UCF and ActivityNet) validate that the proposed AVMST outperforms existing state-of-the-art methods by a significant margin. The code and pre-trained models are available at https://github.com/liwr-hit/ICME23_AVMST. Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Hengyu Man, Xiaopeng Fan 0001 |
ICME | 1 |
| 2023 | Reservoir Computing Transformer for Image-Text RetrievalabstractAlthough the attention mechanism in transformers has proven successful in image-text retrieval tasks, most transformer models suffer from a large number of parameters. Inspired by brain circuits that process information with recurrent connected neurons, we propose a novel Reservoir Computing Transformer Reasoning Network (RCTRN) for image-text retrieval. The proposed RCTRN employs a two-step strategy to focus on feature representation and data distribution of different modalities respectively. Specifically, we send visual and textual features through a unified meshed reasoning module, which encodes multi-level feature relationships with prior knowledge and aggregates the complementary outputs in a more effective way. The reservoir reasoning network is proposed to optimize memory connections between features at different stages and address the data distribution mismatch problem introduced by the unified scheme. To investigate the significance of the low power dissipation and low bandwidth characteristics of RRN in practical scenarios, we deployed the model in the wireless transmission system, demonstrating that RRN's optimization of data structures also has a certain robustness against channel noise. Extensive experiments on two benchmark datasets, Flickr30K and MS-COCO, demonstrate the superiority of RCTRN in terms of performance and low-power dissipation compared to state-of-the-art baselines. Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Penghong Wang, Jinqiao Shi, Xiaopeng Fan 0001 |
ACM Multimedia | 1 |
| 2023 | Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot LearningabstractAudio-visual zero-shot learning (ZSL) has attracted board attention, as it could classify video data from classes that are not observed during training. However, most of the existing methods are restricted to background scene bias and fewer motion details by employing a single-stream network to process scenes and motion information as a unified entity. In this paper, we address this challenge by proposing a novel dual-stream architecture Motion-Decoupled Spiking Transformer (MDFT) to explicitly decouple the contextual semantic information and highly sparsity dynamic motion information. Specifically, The Recurrent Joint Learning Unit (RJLU) could extract contextual semantic information effectively and understand the environment in which actions occur by capturing joint knowledge between different modalities. By converting RGB images to events, our approach effectively captures motion information while mitigating the influence of background scene biases, leading to more accurate classification results. We utilize the inherent strengths of Spiking Neural Networks (SNNs) to process highly sparsity event data efficiently. Additionally, we introduce a Discrepancy Analysis Block (DAB) to model the audio motion features. To enhance the efficiency of SNNs in extracting dynamic temporal and motion information, we dynamically adjust the threshold of Leaky Integrate-and-Fire (LIF) neurons based on the statistical cues of global motion and contextual semantic information. Our experiments demonstrate the effectiveness of MDFT, which consistently outperforms state-of-the-art methods across mainstream benchmarks. Moreover, we find that motion information serves as a powerful regularization for video networks, where using it improves the accuracy of HM and ZSL by 19.1% and 38.4%, respectively. Wenrui Li 0001, Xi-Le Zhao, Zhengyu Ma, Xiaopeng Fan 0001, Yonghong Tian 0001 |
ACM Multimedia | 1 |
| 2023 | The Style Transformer With Common Knowledge Optimization for Image-Text RetrievalabstractImage-text retrieval which associates different modalities has drawn broad attention due to its excellent research value and broad real-world application. However, most of the existing methods haven't taken the high-level semantic relationships (“style embedding”) and common knowledge from multi-modalities into full consideration. To this end, we introduce a novel style transformer network with common knowledge optimization (CKSTN) for image-text retrieval. The main module is the common knowledge adaptor (CKA) with both the style embedding extractor (SEE) and the common knowledge optimization (CKO) modules. Specifically, the SEE uses the sequential update strategy to effectively connect the features of different stages in SEE. The CKO module is introduced to dynamically capture the latent concepts of common knowledge from different modalities. Besides, to get generalized temporal common knowledge, we propose a sequential update strategy to effectively integrate the features of different layers in SEE with previous common feature units. CKSTN demonstrates the superiorities of the state-of-the-art methods in image-text retrieval on MSCOCO and Flickr30 K datasets. Moreover, CKSTN is constructed based on the lightweight transformer which is more convenient and practical for the application of real scenes, due to the better performance and lower parameters. Wenrui Li 0001, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2023 | Neuron-Based Spiking Transmission and Reasoning Network for Robust Image-Text RetrievalabstractMost of the image-text retrieval methods carry out accurate results using fine-grained features for feature alignment. However, extracting the robustness features while maintaining the retrieval accuracy in wireless communication is still a challenge, especially with channel noises and limited transmission bandwidth. Inspired by spike signals of neurons in the human brain, we propose the neuron-based spiking transmission and reasoning network (NSTRN). In this way, the features are compressed into compacted efficient representations. In NSTRN, we construct the feature sender based on spiking activation function to selectively encode only important information in images and sentences into binary codes, and reduce the transmission cost. Moreover, the feature receiver is designed as a recurrent architecture and applies both temporal attention and global attention blocks to memorize long-term information. Finally, to compensate for the loss of visual concepts in transmission, we use the global textual features as coefficients to guide the formation of visual features in the training stage. The traditional CNN-based joint source-channel coding model outputs float-point encoded features, which requires additional quantization steps to convert features into binary bitstreams in the practical wireless communication system. Instead, the spiking neural networks (SNNs) directly use binary spike trains to reduce the computation complexity caused by the quantization steps. More importantly, SNNs can naturally encode the asynchronous event streams and inhibit the discrete noisy events to extract robust information. Even with binary bitstreams, NSTRN shows effectiveness compared with the state-of-the-art image-text retrieval methods. In the wireless communication scenario, NSTRN not only reduces the transmission bandwidth but also alleviates the “cliff effect” to a certain extent in the traditional separate encoding methods. To the best of our knowledge, this is the first work using SNNs on robust image-text retrieval. Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Image-Text Alignment and Retrieval Using Light-Weight TransformerabstractWith the increasing demand for multi-media data retrieval in different modalities, cross-modal retrieval algorithms based on deep learning are constantly updated. However, most of them have trouble with large model parameters and insufficient intrinsic nature between different modalities. We proposed a Light-weight Transformer Alignment Network (LTAN), which adopts the current mainstream visual and textual feature extraction methods. With convolutional neural network combined with light-weight transformer architecture and fully connected neural network, LTAN improves the generalization ability of the model while maintaining high performance. In order to extract visual features that lay emphasis on global details, enhancement paths are constructed to fuse precise location signals stored in low-level features with semantic information extracted from high-level to improve the model retrieval accuracy. It obtains the state-of-the-art results on image and sentence retrieval on MS-COCO and Flickr30k datasets. On the MS-COCO 1K test set, our model obtains an improvement of 3.9% and 2.5% respectively for the image and sentence retrieval tasks on the Recall@1 metric. The size of our model is 15% smaller than models using standard transformer as backbone. Wenrui Li 0001, Xiaopeng Fan 0001 |
ICASSP | 1 |
| 2019 | Non-orthogonal approximate joint diagonalization of non-Hermitian matrices in the least-squares sense
Jifei Miao, Guanghui Cheng, Wenrui Li 0001 |
Neurocomputing | 3 |