Yufei Guo 0001

dblp:23/2981-1 · DBLP profile ↗
← Back
34ranked-venue papers
16as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 13 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 20 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hilbert Curve-Encoded Rotation-Equivariant Oriented Object Detector with Locality-Preserving Spatial Mapping
abstract
Arbitrary-Oriented Object Detection (AOOD) has found broad applications in embodied intelligence, autonomous driving, and satellite remote sensing. However, current AOOD frameworks face challenges in ineffective feature extraction and orientation regression inaccuracy. Inspired by Hilbert curve's intrinsic locality-preserving property, we propose a flexible Hilbert curve-Encoded Rotation-Equivariant Oriented Object Detector (HERO-Det). Our innovations include: (i) a novel Hilbert curve traversal convolution paradigm with a dimensionality reduction scheme, which employs locality-preserving spatial filling curves for feature transformation, (ii) a Hilbert pyramid transformer enabling hierarchical construction of multi-scale feature sequences through space-folding operations, as well as (iii) an orientation-adaptive prediction head that decouples rotation-equivariant regression features from invariant classification cues to resolve orientation regression dilemmas in two-stage detectors. Extensive experiments show HERO-Det achieves state-of-the-art performance on AOOD benchmarks, with mAP of 79.56%, 90.64%, 90.10%, and 80.47% on DOTA, HRSC2016, SSDD, and HRSID, respectively. Performance gains in cross-task validation further demonstrate the versatility of our method to diverse vision tasks, such as medical image segmentation and 3D object detection.
Qi Ming, Liuqian Wang, Ziyi Teng, Xiaoxi Hu, Yufei Guo 0001
AAAI10
2026 STFD-SNN: A Physics-Constrained Spiking Neural Network Framework for Maritime Radio Environment Map Reconstruction
abstract
The escalating disparity between the supply and demand of maritime radio spectrum resources necessitates the construction of high-fidelity Radio Environment Maps (REM) for effective spectrum situational awareness and dynamic management. However, this task is severely hampered by distinctive maritime challenges, including extreme data sparsity, highly dynamic electromagnetic propagation characteristics, and complex spatio-temporal correlations, which significantly degrade conventional terrestrial REM reconstruction methods. To overcome these limitations, this paper proposes a hierarchical REM reconstruction framework that synergistically integrates Spiking Neural Networks with physical constraints. Our contributions are threefold. First, we devise an adaptive Unmanned Aerial Vehicle sampling strategy based on a refined Ant Colony Optimization algorithm, incorporating a hierarchical priority decision mechanism and joint heuristic function to improve data collection efficiency under sparse sampling. Second, we architect a Frequency-Spatio-Temporal Attention (FSTA) -enhanced Spiking Neural Network (SNN) model that captures spatio-temporal dynamics from sparse observations for high-precision 2D REM completion. Third, we introduce a physics-guided knowledge distillation paradigm that embeds maritime electromagnetic propagation models as multi-stage soft constraints through three coordinated mechanisms, which direct input correction via height-weighted physical deviation terms, and loss-level supervision penalizing physically inconsistent 3D reconstructions. Extensive simulations conducted in a high-fidelity maritime scenario demonstrate, which is constructed using real geographic environments, GMTED digital elevation data, and representative meteorological conditions. Our framework outperforms tensor completion U-Net and PINN baseline across various sampling rates. Notably, at sampling rates of 20%, 50%, and 80%, the proposed framework consistently attains superior Root Mean Square Error (RMSE) and Normalized Mean Square Error (NMSE), with the Physical Residual Metric (PRM) further serving as a diagnostic indicator confirming internalization of physical priors.
Liu Yi, Youchen Fan, Yufei Guo 0001, You Fu, Shengliang Fang, Qichen Wang 0014
IEEE Internet Things J.3
2025 Improving Transformer Based Line Segment Detection with Matched Predicting and Re-ranking
abstract
Classical Transformer-based line segment detection methods have delivered impressive results. However, we observe that some accurately detected line segments are assigned low confidence scores during prediction, causing them to be ranked lower and potentially suppressed. Additionally, these models often require prolonged training periods to achieve strong performance, largely due to the necessity of bipartite matching. In this paper, we introduce RANK-LETR, a novel Transformer-based line segment detection method. Our approach leverages learnable geometric information to refine the ranking of predicted line segments by enhancing the confidence scores of high-quality predictions in a posterior verification step. We also propose a new line segment proposal method, wherein the feature point nearest to the centroid of the line segment directly predicts the location, significantly improving training efficiency and stability. Moreover, we introduce a line segment ranking loss to stabilize rankings during training, thereby enhancing the generalization capability of the model. Experimental results demonstrate that our method outperforms other Transformer-based and CNN-based approaches in prediction accuracy while requiring fewer training epochs than previous Transformer-based models.
Shi Peng, Baojie Tian, Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001
AAAI4
2025 Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for Transformer
abstract
Transformers have demonstrated outstanding performance across a wide range of tasks, owing to their self-attention mechanism, but they are highly energy-consuming. Spiking Neural Networks have emerged as a promising energy-efficient alternative to traditional Artificial Neural Networks, leveraging event-driven computation and binary spikes for information transfer. The combination of Transformers’ capabilities with the energy efficiency of SNNs offers a compelling opportunity. This paper addresses the challenge of adapting the self-attention mechanism of Transformers to the spiking paradigm by introducing a novel approach: Accurate Addition-Only Spiking Self-Attention (A2OS2A). Unlike existing methods that rely solely on binary spiking neurons for all components of the self-attention mechanism, our approach integrates binary, ReLU, and ternary spiking neurons. This hybrid strategy significantly improves accuracy while preserving non-multiplicative computations. Moreover, our method eliminates the need for softmax and scaling operations. Extensive experiments show that the A2OS2A-based Spiking Transformer outperforms existing SNN-based Transformers on several datasets, even achieving an accuracy of 78.66% on ImageNet-1K. Our work represents a significant advancement in SNN-based Transformer models, offering a more accurate and efficient solution for real-world applications.
Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Yuhan Zhang 0006, Zhe Ma 0001
CVPR1
2025 Synergistic Prompting for Robust Visual Recognition with Missing Modalities
abstract
Large-scale multi-modal models have demonstrated remarkable performance across various visual recognition tasks by leveraging extensive paired multi-modal training data. However, in real-world applications, the presence of missing or incomplete modality inputs often leads to significant performance degradation. Recent research has focused on prompt-based strategies to tackle this issue; however, existing methods are hindered by two major limitations: (1) static prompts lack the flexibility to adapt to varying missing-data conditions, and (2) basic prompt-tuning methods struggle to ensure reliable performance when critical modalities are missing.To address these challenges, we propose a novel Synergistic Prompting (SyP) framework for robust visual recognition with missing modalities. The proposed SyP introduces two key innovations: (I) a Dynamic Adapter, which computes adaptive scaling factors to dynamically generate prompts, replacing static parameters for flexible multi-modal adaptation, and (II) a Synergistic Prompting Strategy, which combines static and dynamic prompts to balance information across modalities, ensuring robust reasoning even when key modalities are missing. The proposed SyP achieves significant performance improvements over existing approaches across three widely-used visual recognition datasets, demonstrating robustness under diverse missing rates and conditions. Extensive experiments and ablation studies validate its effectiveness in handling missing modalities, highlighting its superior adaptability and reliability.
Luanyuan Dai, Qika Lin, Yunfeng Diao, Guangyin Jin, Yufei Guo 0001, Jing Zhang 0037, Xiaoshuai Hao
ICCV6
2025 ReverB-SNN: Reversing Bit of the Weight and Activation for Spiking Neural Networks
abstract
The Spiking Neural Network (SNN), a biologically inspired neural network infrastructure, has garnered significant attention recently. SNNs utilize binary spike activations for efficient information transmission, replacing multiplications with additions, thereby enhancing energy efficiency. However, binary spike activation maps often fail to capture sufficient data information, resulting in reduced accuracy. To address this challenge, we advocate reversing the bit of the weight and activation, called \textbf{ReverB}, inspired by recent findings that highlight greater accuracy degradation from quantizing activations compared to weights. Specifically, our method employs real-valued spike activations alongside binary weights in SNNs. This preserves the event-driven and multiplication-free advantages of standard SNNs while enhancing the information capacity of activations. Additionally, we introduce a trainable factor within binary weights to adaptively learn suitable weight amplitudes during training, thereby increasing network capacity. To maintain efficiency akin to vanilla \textbf{ReverB}, our trainable binary weight SNNs are converted back to standard form using a re-parameterization technique during inference. Extensive experiments across various network architectures and datasets, both static and dynamic, demonstrate that our approach consistently outperforms state-of-the-art methods.
Yufei Guo 0001, Yuhan Zhang 0006, Jie Zhou 0001, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Zhe Ma 0001
ICML1
2025 RANK++LETR: Learn to Rank and Optimize Candidates for Line Segment Detection
abstract
It is observed that the confidence score may fail to reflect the predicting quality accurately in previous proposal-based line segment detection methods, since the scores and the line locations are predicted simultaneously. We find that the line segment detection performance can be further improved by learning-based line candidate ranking and optimizing strategy. To this end, we build a novel end-to-end line detecting model named RANK++LETR upon deformable DETR architecture, where the encoder is used to select the line candidates while the decoder is applied to rank and optimize these candidates. We design line-aware deformable attention (LADA) module in which attention positions are distributed in a long narrow area and can align well with the elongated geometry of line segments. Moreover, we innovatively apply ranking-based supervision in line segment detection task with the design of contiguous labels according to the detection quality. Experimental results demonstrate that our method outperforms previous SOTA methods in prediction accuracy and gets faster inferring speed than other Transformer-based methods.
Baojie Tian, Yufei Guo 0001, Zhe Ma 0001
NeurIPS3
2025 Spik-NeRF: Spiking Neural Networks for Neural Radiance Fields
abstract
Spiking Neural Networks (SNNs), as a biologically inspired neural network architecture, have garnered significant attention due to their exceptional energy efficiency and increasing potential for various applications. In this work, we extend the use of SNNs to neural rendering tasks and introduce Spik-NeRF (Spiking Neural Radiance Fields). We observe that the binary spike activation map of traditional SNNs lacks sufficient information capacity, leading to information loss and a subsequent decline in the performance of spiking neural rendering models. To address this limitation, we propose the use of ternary spike neurons, which enhance the information-carrying capacity in the spiking neural rendering model. With ternary spike neurons, Spik-NeRF achieves performance that is on par with, or nearly identical to, traditional ANN-based rendering models. Additionally, we present a re-parameterization technique for inference that allows Spik-NeRF with ternary spike neurons to retain the event-driven, multiplication-free advantages typical of binary spike neurons. Furthermore, to further boost the performance of Spik-NeRF, we employ a distillation method, using an ANN-based NeRF to guide the training of our Spik-NeRF model, which is more compatible with the ternary neurons compared to the standard binary neurons. We evaluate Spik-NeRF on both realistic and synthetic scenes, and the experimental results demonstrate that Spik-NeRF achieves rendering performance comparable to ANN-based NeRF models.
Qinlong Lan, Yitian Wu, Z. Jane Wang 0001, Wanhua Li 0004, Yufei Guo 0001
NeurIPS8
2025 UP-Diff: Latent Diffusion Model for Remote Sensing Urban Prediction
abstract
Remote sensing (RS) technology has become essential for monitoring urban development, including applications like population growth analysis, transportation congestion forecasting, and climate change detection (CD). However, its potential for future urban planning (UP), particularly in predicting future urban layouts remains largely unexplored. This study introduces UP-Diff, a novel method leveraging generative models for UP, to address this gap. UP-Diff leverages information from current urban layouts and planned change maps to predict future urban configurations. Key challenges addressed include the integration of urban layouts and change maps into latent diffusion model (LDM) through careful architecture improvements and the mitigation of limited training data by employing a pretrained stable diffusion (SD) model with fixed weights, trainable ConvNeXt, and trainable cross-attention layers. Our method significantly streamlines the UP process by automating layout predictions, thus reducing the time and effort required compared to traditional manual methods. Comprehensive evaluations on the learning, vision, and RS dataset (LEVIR-CD) and Sun Yat-Sen University dataset (SYSU-CD) validate that UP-Diff achieves high-fidelity predictions of future urban layouts, demonstrating its effectiveness and potential for advancing RS-based UP methodologies. Our code and model weights are available athttps://github.com/zeyuwang-zju/UP-Diff.
Zeyu Wang 0010, Zecheng Hao, Yuhan Zhang 0006, Yuchao Feng, Yufei Guo 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 PT-BitNet: Scaling up the 1-Bit large language model with post-training quantization
Yufei Guo 0001, Zecheng Hao, Jiahang Shao, Jie Zhou 0001, Xiaode Liu, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Zhe Ma 0001
Neural Networks1
2025 All-in-Focus Imaging From Events With Occlusions
abstract
Event-based Synthetic Aperture Imaging (E-SAI) extends the SAI technique to observe targets behind extremely dense occlusions. Existing approaches remain confined to the de-occlusion of a specific depth plane, i.e., single depth in focus, unable to be applied to observe occluded targets with varying depths due to the decreased focus range. To achieve All-in-Focus E-SAI, i.e., recovering the occlusion-free image of all depth planes, the depth information behind the occlusions should be given to ensure accurate event refocusing. In this paper, we first prove the feasibility of predicting the depth map from captured events in the presence of dense occlusions. Then, we propose the ESAI-AF network, which consists of a Depth Estimation Module (DEM) designed to estimate the depth information from multi-view events and an Image Enhancement Module (IEM) designed to reconstruct high-quality occlusion-free images from the refocused events. We employ only multi-view occlusion-free images as supervised signals for end-to-end training of the above modules. Extensive experiments have shown that the proposed method can effectively perform All-in-Focus image reconstruction of occluded multi-depth targets and achieves superior performance to existing methods.
Lixuan Wei, Yufei Guo 0001, Lei Yu 0006
IEEE Trans. Multim.3
2025 Visual-Semantic Graph Matching Net for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to leverage additional semantic information to recognize unseen classes. To transfer knowledge from seen to unseen classes, most ZSL methods often learn a shared embedding space by simply aligning visual embeddings with semantic prototypes. However, methods trained under this paradigm often struggle to learn robust embedding space because they align the two modalities in an isolated manner among classes, which ignore the crucial class relationship during the alignment process. To address the aforementioned challenges, this article proposes a visual-semantic graph matching net (VSGMN), which leverages semantic relationships among classes to aid in visual-semantic embedding. VSGMN uses a graph build net (GBN) and a graph matching net (GMN) to achieve two-stage visual-semantic alignment. Specifically, GBN first uses an embedding-based approach to build visual and semantic graphs in the semantic space and align the embedding with its prototype for first-stage alignment. In addition, to supplement unseen class relationships in these graphs, GBN also builds the unseen class nodes based on semantic relationships. In the second stage, GMN continuously integrates neighbor and cross-graph information into the constructed graph nodes and aligns the node relationships between the two graphs under the class relationship constraint. Extensive experiments on three benchmark datasets demonstrate that VSGMN achieves superior performance in both conventional and generalized ZSL (GZSL) scenarios. The implementation of our VSGMN and experimental results are available at github: https://github.com/dbwfd/VSGMN.
Bowen Duan 0001, Shiming Chen 0002, Yufei Guo 0001, Guosen Xie, Weiping Ding 0001, Yisong Wang 0004
IEEE Trans. Neural Networks Learn. Syst.3
2024 Ternary Spike: Learning Ternary Spikes for Spiking Neural Networks
abstract
The Spiking Neural Network (SNN), as one of the biologically inspired neural network infrastructures, has drawn increasing attention recently. It adopts binary spike activations to transmit information, thus the multiplications of activations and weights can be substituted by additions, which brings high energy efficiency. However, in the paper, we theoretically and experimentally prove that the binary spike activation map cannot carry enough information, thus causing information loss and resulting in accuracy decreasing. To handle the problem, we propose a ternary spike neuron to transmit information. The ternary spike neuron can also enjoy the event-driven and multiplication-free operation advantages of the binary spike neuron but will boost the information capacity. Furthermore, we also embed a trainable factor in the ternary spike neuron to learn the suitable spike amplitude, thus our SNN will adopt different spike amplitudes along layers, which can better suit the phenomenon that the membrane potential distributions are different along layers. To retain the efficiency of the vanilla ternary spike, the trainable ternary spike SNN will be converted to a standard one again via a re-parameterization technique in the inference. Extensive experiments with several popular network structures over static and dynamic datasets show that the ternary spike can consistently outperform state-of-the-art methods. Our code is open-sourced at https://github.com/yfguo91/Ternary-Spike.
Yufei Guo 0001, Yuanpei Chen, Xiaode Liu, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001
AAAI1
2024 End-to-End Real-Time Vanishing Point Detection with Transformer
abstract
In this paper, we propose a novel transformer-based end-to-end real-time vanishing point detection method, which is named Vanishing Point TRansformer (VPTR). The proposed method can directly regress the locations of vanishing points from given images. To achieve this goal, we pose vanishing point detection as a point object detection task on the Gaussian hemisphere with region division. Considering low-level features always provide more geometric information which can contribute to accurate vanishing point prediction, we propose a clear architecture where vanishing point queries in the decoder can directly gather multi-level features from CNN backbone with deformable attention in VPTR. Our method does not rely on line detection or Manhattan world assumption, which makes it more flexible to use. VPTR runs at an inferring speed of 140 FPS on one NVIDIA 3090 card. Experimental results on synthetic and real-world datasets demonstrate that our method can be used in both natural and structural scenes, and is superior to other state-of-the-art methods on the balance of accuracy and efficiency.
Shi Peng, Yufei Guo 0001, Xuhui Huang
AAAI3
2024 Enhancing Representation of Spiking Neural Networks via Similarity-Sensitive Contrastive Learning
abstract
Spiking neural networks (SNNs) have attracted intensive attention as a promising energy-efficient alternative to conventional artificial neural networks (ANNs) recently, which could transmit information in form of binary spikes rather than continuous activations thus the multiplication of activation and weight could be replaced by addition to save energy. However, the binary spike representation form will sacrifice the expression performance of SNNs and lead to accuracy degradation compared with ANNs. Considering improving feature representation is beneficial to training an accurate SNN model, this paper focuses on enhancing the feature representation of the SNN. To this end, we establish a similarity-sensitive contrastive learning framework, where SNN could capture significantly more information from its ANN counterpart to improve representation by Mutual Information (MI) maximization with layer-wise sensitivity to similarity. In specific, it enriches the SNN’s feature representation by pulling the positive pairs of SNN's and ANN's feature representation of each layer from the same input samples closer together while pushing the negative pairs from different samples further apart. Experimental results show that our method consistently outperforms the current state-of-the-art algorithms on both popular non-spiking static and neuromorphic datasets.
Yuhan Zhang 0006, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001
AAAI5
2024 Towards Understanding Factual Knowledge of Large Language Models
abstract
Large language models (LLMs) have recently driven striking performance improvements across a range of natural language processing tasks. The factual knowledge acquired during pretraining and instruction tuning can be useful in various downstream tasks, such as question answering, and language generation. Unlike conventional Knowledge Bases (KBs) that explicitly store factual knowledge, LLMs implicitly store facts in their parameters. Content generated by the LLMs can often exhibit inaccuracies or deviations from the truth, due to facts that can be incorrectly induced or become obsolete over time. To this end, we aim to explore the extent and scope of factual knowledge within LLMs by designing the benchmark Pinocchio. Pinocchio contains 20K diverse factual questions that span different sources, timelines, domains, regions, and languages. Furthermore, we investigate whether LLMs can compose multiple facts, update factual knowledge temporally, reason over multiple pieces of facts, identify subtle factual differences, and resist adversarial examples. Extensive experiments on different sizes and types of LLMs show that existing LLMs still lack factual knowledge and suffer from various spurious correlations. We believe this is a critical bottleneck for realizing trustworthy artificial intelligence. The dataset Pinocchio and our codes are publicly available at: https://github.com/THU-BPM/Pinocchio.
Xuming Hu, Junzhe Chen 0001, Xiaochuan Li 0003, Yufei Guo 0001, Lijie Wen 0001, Philip S. Yu, Zhijiang Guo
ICLR4
2024 EnOF-SNN: Training Accurate Spiking Neural Networks via Enhancing the Output Feature
abstract
Spiking neural networks (SNNs) have gained more and more interest as one of the energy-efficient alternatives of conventional artificial neural networks (ANNs). They exchange 0/1 spikes for processing information, thus most of the multiplications in networks can be replaced by additions. However, binary spike feature maps will limit the expressiveness of the SNN and result in unsatisfactory performance compared with ANNs. It is shown that a rich output feature representation, i.e., the feature vector before classifier) is beneficial to training an accurate model in ANNs for classification. We wonder if it also does for SNNs and how to improve the feature representation of the SNN. To this end, we materialize this idea in two special designed methods for SNNs. First, inspired by some ANN-SNN methods that directly copy-paste the weight parameters from trained ANN with light modification to homogeneous SNN can obtain a well-performed SNN, we use rich information of the weight parameters from the trained ANN counterpart to guide the feature representation learning of the SNN. In particular, we present the SNN's and ANN's feature representation from the same input to ANN's classifier to product SNN's and ANN's outputs respectively and then align the feature with the KL-divergence loss as in knowledge distillation methods, called L_ AF loss. It can be seen as a novel and effective knowledge distillation method specially designed for the SNN that comes from both the knowledge distillation and ANN-SNN methods. Various ablation study shows that the L_AF loss is more powerful than the vanilla knowledge distillation method. Second, we replace the last Leaky Integrate-and-Fire (LIF) activation layer as the ReLU activation layer to generate the output feature, thus a more powerful SNN with full-precision feature representation can be achieved but with only a little extra computation. Experimental results show that our method consistently outperforms the current state-of-the-art algorithms on both popular non-spiking static and neuromorphic datasets. We provide an extremely simple but effective way to train high-accuracy spiking neural networks.
Yufei Guo 0001, Weihang Peng 0001, Xiaode Liu, Yuanpei Chen, Yuhan Zhang 0006, Zhou Jie, Zhe Ma 0001
NeurIPS1
2024 Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks
abstract
The Spiking Neural Network (SNN) is a biologically inspired neural network infrastructure that has recently garnered significant attention. It utilizes binary spike activations to transmit information, thereby replacing multiplications with additions and resulting in high energy efficiency. However, training an SNN directly poses a challenge due to the undefined gradient of the firing spike process. Although prior works have employed various surrogate gradient training methods that use an alternative function to replace the firing process during back-propagation, these approaches ignore an intrinsic problem: gradient vanishing. To address this issue, we propose a shortcut back-propagation method in the paper, which advocates for transmitting the gradient directly from the loss to the shallow layers. This enables us to present the gradient to the shallow layers directly, thereby significantly mitigating the gradient vanishing problem. Additionally, this method does not introduce any burden during the inference phase. To strike a balance between final accuracy and ease of training, we also propose an evolutionary training framework and implement it by inducing a balance coefficient that dynamically changes with the training epoch, which further improves the network's performance. Extensive experiments conducted over static and dynamic datasets using several popular network structures reveal that our method consistently outperforms state-of-the-art methods.
Yufei Guo 0001, Yuanpei Chen, Zecheng Hao, Weihang Peng 0001, Zhou Jie, Yuhan Zhang 0006, Xiaode Liu, Zhe Ma 0001
NeurIPS1
2024 NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN
abstract
Recently, the neuromorphic vision sensor has received more and more interest. However, the neuromorphic data consists of asynchronous event spikes, which makes it difficult to construct a big benchmark to train a power general neural network model, thus limiting the neuromorphic data understanding for “unseen” objects by deep learning. While for the frame image, since the training data can be obtained easily, the zero-shot and few-shot learning for “unseen” task via the large Contrastive Vision-Language Pre-training (CLIP) model, which is pre-trained by large-scale image-text pairs in 2D, have shown inspirational performance. We wonder whether the CLIP could be transferred to neuromorphic data recognition to handle the “unseen” problem. To this end, we materialize this idea with NeuroCLIP in the paper. The NeuroCLIP consists of 2D CLIP and two specially designed modules for neuromorphic data understanding. First, an event-frame module that could convert the event spikes to the sequential frame image with a simple discrimination strategy. Second, an inter-timestep adapter, which is a simple fine-tuned adapter based on a spiking neural network (SNN) for the sequential features coming from the visual encoder of CLIP to improve the few-shot performance. Various experiments on neuromorphic datasets including N-MNIST, CIFAR10-DVS, and ES-ImageNet demonstrate the effectiveness of NeuroCLIP.
Yufei Guo 0001, Yuanpei Chen, Zhe Ma 0001
IEEE Signal Process. Lett.1
2024 Improved Event-Based Image De-Occlusion
abstract
Reconstructing clear scene images in the presence of foreground occlusions remains a formidable challenge for cameras constrained by a single viewpoint. Synthetic aperture imaging (SAI) has emerged as a solution by integrating visual information from multiple viewpoints, thus overcoming occlusions. Event cameras, characterized by their high temporal resolution, exceptional dynamic range, low power consumption, and resistance to motion blur, offer a promising avenue for capturing intricate details of background objects within a limited range of motion. Leveraging these capabilities, several event-camera-based SAI methodologies have been proposed to effectively tackle dense occlusions. Despite these advancements, existing methodologies encounter obstacles stemming from the hybrid model they employ, comprising a spiking neural network (SNN) encoder and a convolutional neural network (CNN) decoder. Challenges include information degradation within the SNN encoder due to the quantization of full-precision data into binary spikes, as well as insufficient training data for the CNN decoder to adequately learn feature extraction. In response, we present an enhanced event-based image de-occlusion approach. Our method introduces a novel full-precision leaky integrate-and-fire (FP-LIF) mechanism to mitigate information loss within the SNN encoder. Additionally, we propose an isomorphic network knowledge distillation method to augment the feature extraction capabilities of the CNN decoder. Experimental results demonstrate the efficacy of our approach in enhancing event-camera-based SAI methodologies.
Yufei Guo 0001, Weihang Peng 0001, Yuanpei Chen, Jie Zhou 0001, Zhe Ma 0001
IEEE Signal Process. Lett.1
2024 Pair-ID: A Dual Modal Framework for Identity Preserving Image Generation
abstract
The acquisition of large-scale paired visible and thermal images is crucial for enhancing face recognition systems, especially in low-light environments where visible spectrum images fail. However, the task is hindered by the scarcity of thermal images and the need for identity consistency during image generation. In this paper, we propose Pair-ID, an innovative framework that addresses these challenges by creating a shared latent space for simultaneous generation of paired visible and thermal images. Pair-ID integrates identity information into text embeddings and employs fixed templates for diverse facial poses, streamlining the customization process and reducing computational demands. The framework's Joint Learner encodes both modalities, facilitating synchronized image generation and preserving facial details. Extensive evaluations show that Pair-ID surpasses current methods in efficiency and performance for paired data generation, making it a promising solution for face recognition under varying lighting conditions.
Yongrong Wu, Zeyu Wang 0010, Xiaode Liu, Yufei Guo 0001
IEEE Signal Process. Lett.5
2024 Event-Based Shutter Unrolling and Motion Deblurring in Dynamic Scenes
abstract
The Rolling Shutter (RS) effect and motion blur are common challenges in images captured by CMOS cameras during dynamic scenes. Inspired by biological vision principles, event cameras capture intensity changes asynchronously with low latency, providing valuable insights into image degradation during exposure. This study addresses the dual challenges of rolling shutter correction and deblurring using event data, merging them into a unified one-stage network. This streamlined approach reduces cumulative errors and inference time compared to traditional two-stage methods. To achieve this, we introduce an Event Representation for Rolling Shutter Deblurring, which explicitly models the conversion relationship between the input RS blurry frame and the latent image using events. To enhance the fusion of image and event information, we present a Time-guided Cross-Modal Attention module. Furthermore, we improve performance by incorporating a Multi-Scale Context-Aware Transformer Block, effectively addressing varying degrees of distortion and blurriness using a multi-scale attention mechanism. Extensive experiments validate that our method outperforms existing state-of-the-art approaches.
Yangguang Wang, Chenxu Jiang, Xu Jia 0012, Yufei Guo 0001, Lei Yu 0006
IEEE Signal Process. Lett.4
2023 Deep Dive into Gradients: Better Optimization for 3D Object Detection with Gradient-Corrected IoU Supervision
abstract
Intersection-over-Union (IoU) is the most popular metric to evaluate regression performance in 3D object detection. Recently, there are also some methods applying IoU to the optimization of 3D bounding box regression. However, we demonstrate through experiments and mathematical proof that the 3D IoU loss suffers from abnormal gradient w.r.t. angular error and object scale, which further leads to slow convergence and suboptimal regression process, respectively. In this paper, we propose a Gradient-Corrected IoU (GCIoU) loss to achieve fast and accurate 3D bounding box regression. Specifically, a gradient correction strategy is designed to endow 3D IoU loss with a reasonable gradient. It ensures that the model converges quickly in the early stage of training, and helps to achieve fine-grained refinement of bounding boxes in the later stage. To solve suboptimal regression of 3D IoU loss for objects at different scales, we introduce a gradient rescaling strategy to adaptively optimize the step size. Finally, we integrate GCIoU Loss into multiple models to achieve stable performance gains and faster model convergence. Experiments on KITTI dataset demonstrate superiority of the proposed method. The code is available at https://github.com/ming71/GCIoU-loss.
Qi Ming, Lingjuan Miao, Zhe Ma 0001, Zhiqiang Zhou 0001, Xuhui Huang, Yuanpei Chen, Yufei Guo 0001
CVPR8
2023 PeakConv: Learning Peak Receptive Field for Radar Semantic Segmentation
abstract
The modern machine learning-based technologies have shown considerable potential in automatic radar scene understanding. Among these efforts, radar semantic segmentation (RSS) can provide more refined and detailed information including the moving objects and background clutters within the effective receptive field of the radar. Motivated by the success of convolutional networks in various visual computing tasks, these networks have also been introduced to solve RSS task. However, neither the regular convolution operation nor the modified ones are specific to interpret radar signals. The receptive fields of existing convolutions are defined by the object presentation in optical signals, but these two signals have different perception mechanisms. In classic radar signal processing, the object signature is detected according to a local peak response, i.e., CFAR detection. Inspired by this idea, we redefine the receptive field of the convolution operation as the peak receptive field (PRF) and propose the peak convolution operation (PeakConv) to learn the object signatures in an end-to-end network. By incorporating the proposed PeakConv layers into the encoders, our RSS network can achieve better segmentation results compared with other SoTA methods on a multi-view real-measured dataset collected from an FMCW radar. Our code for PeakConv is available at https://github.com/zlw9161/PKC.
Liwen Zhang 0001, Youcheng Zhang, Yufei Guo 0001, Yuanpei Chen, Xuhui Huang, Zhe Ma 0001
CVPR4
2023 RMP-Loss: Regularizing Membrane Potential Distribution for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) as one of the biology-inspired models have received much attention recently. It can significantly reduce energy consumption since they quantize the real-valued membrane potentials to 0/1 spikes to transmit information thus the multiplications of activations and weights can be replaced by additions when implemented on hardware. However, this quantization mechanism will inevitably introduce quantization error, thus causing catastrophic information loss. To address the quantization error problem, we propose a regularizing membrane potential loss (RMP-Loss) to adjust the distribution which is directly related to quantization error to a range close to the spikes. Our method is extremely simple to implement and straightforward to train an SNN. Furthermore, it is shown to consistently outperform previous state-of-the-art methods over different network architectures and datasets.
Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Liwen Zhang 0001, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001
ICCV1
2023 Membrane Potential Batch Normalization for Spiking Neural Networks
abstract
As one of the energy-efficient alternatives of conventional neural networks (CNNs), spiking neural networks (SNNs) have gained more and more interest recently. To train the deep models, some effective batch normalization (BN) techniques are proposed in SNNs. All these BNs are suggested to be used after the convolution layer as usually doing in CNNs. However, the spiking neuron is much more complex with the spatio-temporal dynamics. The regulated data flow after the BN layer will be disturbed again by the membrane potential updating operation before the firing function, i.e., the nonlinear activation. Therefore, we advocate adding another BN layer before the firing function to normalize the membrane potential again, called MPBN. To eliminate the induced time cost of MPBN, we also propose a training-inference-decoupled re-parameterization technique to fold the trained MPBN into the firing threshold. With the re-parameterization technique, the MPBN will not introduce any extra time burden in the inference. Furthermore, the MPBN can also adopt the element-wised form, while these BNs after the convolution layer can only use the channel-wised form. Experimental results show that the proposed MPBN performs well on both popular non-spiking static and neuromorphic datasets.
Yufei Guo 0001, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Liwen Zhang 0001, Xuhui Huang, Zhe Ma 0001
ICCV1
2023 Spiking PointNet: Spiking Neural Networks for Point Clouds
abstract
Recently, Spiking Neural Networks (SNNs), enjoying extreme energy efficiency, have drawn much research attention on 2D visual recognition and shown gradually increasing application potential. However, it still remains underexplored whether SNNs can be generalized to 3D recognition. To this end, we present Spiking PointNet in the paper, the first spiking neural model for efficient deep learning on point clouds. We discover that the two huge obstacles limiting the application of SNNs in point clouds are: the intrinsic optimization obstacle of SNNs that impedes the training of a big spiking model with large time steps, and the expensive memory and computation cost of PointNet that makes training a big spiking point model unrealistic. To solve the problems simultaneously, we present a trained-less but learning-more paradigm for Spiking PointNet with theoretical justifications and in-depth experimental analysis. In specific, our Spiking PointNet is trained with only a single time step but can obtain better performance with multiple time steps inference, compared to the one trained directly with multiple time steps. We conduct various experiments on ModelNet10, ModelNet40 to demonstrate the effectiveness of Sipiking PointNet. Notably, our Spiking PointNet even can outperform its ANN counterpart, which is rare in the SNN field thus providing a potential research direction for the following work. Moreover, Spiking PointNet shows impressive speedup and storage saving in the training phase. Our code is open-sourced at https://github.com/DayongRen/Spiking-PointNet.
Dayong Ren, Zhe Ma 0001, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Yuhan Zhang 0006, Yufei Guo 0001
NeurIPS7
2023 Joint A-SNN: Joint training of artificial and spiking neural networks via self-Distillation and weight factorization
Yufei Guo 0001, Weihang Peng 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Xuhui Huang, Zhe Ma 0001
Pattern Recognit.1
2022 RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks
abstract
The brain-inspired and event-driven Spiking Neural Network (SNN) aiming at mimicking the synaptic activity of biological neurons has received increasing attention. It transmits binary spike signals between network units when the membrane potential exceeds the firing threshold. This biomimetic mechanism of SNN appears energy-efficiency with its power sparsity and asynchronous operations on spike events. Unfortunately, with the propagation of binary spikes, the distribution of membrane potential will shift, leading to degeneration, saturation, and gradient mismatch problems, which would be disadvantageous to the network optimization and convergence. Such undesired shifts would prevent the SNN from performing well and going deep. To tackle these problems, we attempt to rectify the membrane potential distribution (MPD) by designing a novel distribution loss, MPD-Loss, which can explicitly penalize the un-desired shifts without introducing any additional operations in the inference phase. Moreover, the proposed method can also mitigate the quantization error in SNNs, which is usually ignored in other works. Experimental results demonstrate that the proposed method can directly train a deeper, larger, and better-performing SNN within fewer timesteps.
Yufei Guo 0001, Xinyi Tong 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Zhe Ma 0001, Xuhui Huang
CVPR1
2022 Reducing Information Loss for Spiking Neural Networks
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, YingLei Wang, Xiaode Liu, Xinyi Tong 0001, Yuanyuan Ou, Xuhui Huang, Zhe Ma 0001
ECCV (11)1
2022 Real Spike: Learning Real-Valued Spikes for Spiking Neural Networks
Yufei Guo 0001, Liwen Zhang 0001, Yuanpei Chen, Xinyi Tong 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001
ECCV (12)1
2022 IM-Loss: Information Maximization Loss for Spiking Neural Networks
abstract
Spiking Neural Network (SNN), recognized as a type of biologically plausible architecture, has recently drawn much research attention. It transmits information by $0/1$ spikes. This bio-mimetic mechanism of SNN demonstrates extreme energy efficiency since it avoids any multiplications on neuromorphic hardware. However, the forward-passing $0/1$ spike quantization will cause information loss and accuracy degradation. To deal with this problem, the Information maximization loss (IM-Loss) that aims at maximizing the information flow in the SNN is proposed in the paper. The IM-Loss not only enhances the information expressiveness of an SNN directly but also plays a part of the role of normalization without introducing any additional operations (\textit{e.g.}, bias and scaling) in the inference phase. Additionally, we introduce a novel differentiable spike activity estimation, Evolutionary Surrogate Gradients (ESG) in SNNs. By appointing automatic evolvable surrogate gradients for spike activity function, ESG can ensure sufficient model updates at the beginning and accurate gradients at the end of the training, resulting in both easy convergence and high task performance. Experimental results on both popular non-spiking static and neuromorphic datasets show that the SNN models trained by our method outperform the current state-of-the-art algorithms.
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001
NeurIPS1
2021 An Improved Advancing-front-Delaunay Method for Triangular Mesh Generation
Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001, Yongqing Hai, Rongli Zhao, Kewu Sun
CGI1
2021 Differentiable Spike: Rethinking Gradient-Descent for Training Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have emerged as a biology-inspired method mimicking the spiking nature of brain neurons. This bio-mimicry derives SNNs' energy efficiency of inference on neuromorphic hardware. However, it also causes an intrinsic disadvantage in training high-performing SNNs from scratch since the discrete spike prohibits the gradient calculation. To overcome this issue, the surrogate gradient (SG) approach has been proposed as a continuous relaxation. Yet the heuristic choice of SG leaves it vacant how the SG benefits the SNN training. In this work, we first theoretically study the gradient descent problem in SNN training and introduce finite difference gradient to quantitatively analyze the training behavior of SNN. Based on the introduced finite difference gradient, we propose a new family of Differentiable Spike (Dspike) functions that can adaptively evolve during training to find the optimal shape and smoothness for gradient estimation. Extensive experiments over several popular network structures show that training SNN with Dspike consistently outperforms the state-of-the-art training methods. For example, on the CIFAR10-DVS classification task, we can train a spiking ResNet-18 and achieve 75.4% top-1 accuracy with 10 time steps.
Yuhang Li 0001, Yufei Guo 0001, Shanghang Zhang, Shikuang Deng, Yongqing Hai, Shi Gu
NeurIPS2