Qinghai Guo

dblp:12/8502 · DBLP profile ↗
← Back
30ranked-venue papers
0as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Hessian-Aware Zeroth-Order Optimization for Quantized Large Language Models
Zhangchi Wang, Yidu Wu, Yijin Huang, Pujin Cheng, Qinghai Guo, Xiaoying Tang 0001
ICIC (5)5
2026 Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs
Yidu Wu, Kejie Zhao, Zhangchi Wang, Qinghai Guo, Xiaoying Tang 0001
ICIC (16)5
2026 Decoupled spatial-temporal predicting model for weakly supervised action localization
Guiqin Wang, Peng Zhao 0001, Shusen Yang, Qinghai Guo
Knowl. Based Syst.7
2026 Implicit hierarchical temporal-spatial residual model for long-term video prediction
Guiqin Wang, Peng Zhao 0001, Haoran Guo, Cong Zhao 0001, Qinghai Guo, Shusen Yang
Neural Networks6
2026 DHPT: Dual-Modality Heterogeneous Prompt Tuning for Online Test-Time Adaption in Vision-Language Models
abstract
Test-Time Adaptation (TTA) has recently emerged as a promising research direction, enabling vision-language models (VLMs) to adapt to unlabeled test data in zero-shot settings. Among TTA approaches, test-time prompt tuning has shown great potential for enhancing the practical applicability of VLMs. However, existing methods typically either focus on adapting a single modality or apply uniform optimization to both modalities, without explicitly defining modality-specific optimization objectives. Such a one-size-fits-all strategy often results in suboptimal performance under test-time conditions. To address this limitation, we propose Dual-modality Heterogeneous Prompt Tuning (DHPT), a novel framework designed to simultaneously capture fine-grained textual semantics and alleviate domain shift noise in the visual modality. Specifically, we leverage a large language model to provide textual cognition guidance for the text encoder, while on the vision side, we develop a lightweight calibration module that adaptively mitigates domain shift noise across different scales. Furthermore, we introduce a cluster-tight optimization objective that enhances the stability and generalizability of prompt tuning under distribution shifts. Extensive experiments conducted on 11 benchmark datasets demonstrate that DHPT consistently and significantly outperforms existing TTA methods for VLMs.
Guiqin Wang, Peng Zhao 0001, Haoran Guo, Shusen Yang, Qinghai Guo
IEEE Trans. Circuits Syst. Video Technol.7
2025 SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
abstract
Known as low energy consumption networks, spiking neural networks (SNNs) have gained a lot of attention within the past decades. While SNNs are increasing competitive with artificial neural networks (ANNs) for vision tasks, they are rarely used for long sequence tasks, despite their intrinsic temporal dynamics. In this work, we develop spiking state space models (SpikingSSMs) for long sequence learning by leveraging on the sequence learning abilities of state space models (SSMs). Inspired by dendritic neuron structure, we hierarchically integrate neuronal dynamics with the original SSM block, meanwhile realizing sparse synaptic computation. Furthermore, to solve the conflict of event-driven neuronal dynamics with parallel computing, we propose a light-weight surrogate dynamic network which accurately predicts the after-reset membrane potential and compatible to learnable thresholds, enabling orders of acceleration in training speed compared with conventional iterative methods. On the long range arena benchmark task, SpikingSSM achieves competitive performance to state-of-the-art SSMs meanwhile realizing on average 90% of network sparsity. On language modeling, our network significantly surpasses existing spiking large language models (spikingLLMs) on the WikiText-103 dataset with only a third of the model size, demonstrating its potential as backbone architecture for low computation cost LLMs.
Shuaijie Shen, Renzhuo Huang, Yan Zhong 0001, Qinghai Guo, Zhichao Lu, Jianguo Zhang 0001, Luziwei Leng
AAAI5
2025 VISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural Network
abstract
Visual streaming perception (VSP) involves online intelligent processing of sequential frames captured by vision sensors, enabling real-time decision-making in applications such as autonomous driving, UAVs, and AR/VR. However, the computational efficiency of VSP on edge devices remains a challenge due to power constraints and the under-utilization of temporal dependencies between frames. While spiking neural networks (SNNs) offer biologically inspired event-driven processing with potential energy benefits, their practical advantage over artificial neural networks (ANNs) for VSP tasks remains unproven. In this work, we introduce a novel framework, ViStream, which leverages the Law of Charge Conservation (LoCC) property in ST-BIF neurons and a differential encoding (DiffEncode) scheme to optimize SNN inference for VSP. By encoding temporal differences between neighboring frames and eliminating frequent membrane resets, ViStream achieves significant computational reduction while maintaining accuracy equivalent to its ANN counterpart. We provide theoretical proofs of equivalence and validate ViStream across diverse VSP tasks, including object detection, tracking, and segmentation, demonstrating substantial energy savings without compromising performance. ViStream is publicly available at: https://github.com/Intelligent-Computing-Research-Group/ViStream
Kang You, Ziling Wei, Qinghai Guo, Zhezhi He
CVPR5
2025 Explicit Mutual Information Maximization for Self-Supervised Learning
abstract
Recently, self-supervised learning (SSL) has been extensively studied. Theoretically, mutual information maximization (MIM) is an optimal criterion for SSL, with a strong theoretical foundation in information theory. However, it is difficult to directly apply MIM in SSL since the data distribution is not analytically available in applications. In practice, many existing methods can be viewed as approximate implementations of the MIM criterion. This work shows that, based on the invariance property of MI, explicit MI maximization can be applied to SSL under a generic distribution assumption, i.e., a relaxed condition of the data distribution. We further illustrate this by analyzing the generalized Gaussian distribution. Based on this result, we derive a loss function based on the MIM criterion using only second-order statistics. We implement the new loss for SSL and demonstrate its effectiveness via extensive experiments.
Lele Chang, Qinghai Guo, Fei Wen 0005
ICASSP3
2025 Threshold Modulation for Online Test-Time Adaptation of Spiking Neural Networks
abstract
Recently, spiking neural networks (SNNs), deployed on neuromorphic chips, provide highly efficient solutions on edge devices in different scenarios. However, their ability to adapt to distribution shifts after deployment has become a crucial challenge. Online test-time adaptation (OTTA) offers a promising solution by enabling models to dynamically adjust to new data distributions without requiring source data or labeled target samples. Nevertheless, existing OTTA methods are largely designed for traditional artificial neural networks and are not well-suited for SNNs. To address this gap, we propose a low-power, neuromorphic chip-friendly online test-time adaptation framework, aiming to enhance model generalization under distribution shifts. The proposed approach is called Threshold Modulation (TM), which dynamically adjusts the firing threshold through neuronal dynamics-inspired normalization, being more compatible with neuromorphic hardware. Experimental results on benchmark datasets demonstrate the effectiveness of this method in improving the robustness of SNNs against distribution shifts while maintaining low computational cost. The proposed method offers a practical solution for online test-time adaptation of SNNs, providing inspiration for the design of future neuromorphic chips. The demo code is available at github.com/NneurotransmitterR/TM-OTTA-SNN.
Kejie Zhao, Wenjia Hua, Aiersi Tuerhong, Luziwei Leng, Yuxin Ma 0001, Qinghai Guo
IJCNN6
2025 Accurate and Efficient Event-Based Semantic Segmentation Using Adaptive Spiking Encoder-Decoder Network
abstract
Spiking neural networks (SNNs), known for their low-power, event-driven computation, and intrinsic temporal dynamics, are emerging as promising solutions for processing dynamic, asynchronous signals from event-based sensors. Despite their potential, SNNs face challenges in training and architectural design, resulting in limited performance in challenging event-based dense prediction tasks compared with artificial neural networks (ANNs). In this work, we develop an efficient spiking encoder-decoder network (SpikingEDN) for large-scale event-based semantic segmentation (EbSS) tasks. To enhance the learning efficiency from dynamic event streams, we harness the adaptive threshold which improves network accuracy, sparsity, and robustness in streaming inference. Moreover, we develop a dual-path spiking spatially adaptive modulation (SSAM) module, which is specifically tailored to enhance the representation of sparse events and multimodal inputs, thereby considerably improving network performance. Our SpikingEDN attains a mean intersection over union (MIoU) of 72.57% on the DDD17 dataset and 58.32% on the larger DSEC-Semantic dataset, showing competitive results to the state-of-the-art ANNs while requiring substantially fewer computational resources. Our results shed light on the untapped potential of SNNs in event-based vision applications. The source codes are publicly available at https://github.com/EMI-Group/spikingedn.
Luziwei Leng, Kaiwei Che, Qinghai Guo, Jianxing Liao, Ran Cheng 0004
IEEE Trans. Neural Networks Learn. Syst.6
2024 Fusion Side Tuning: A Parameter and Memory Efficient Fine-tuning Method for High-resolution Medical Image Classification
abstract
Parameter-efficient fine-tuning (PEFT) has been proposed as a cost-effective approach for transferring large-scale pre-trained models (LPMs) to downstream tasks, mitigating the high costs associated with updating all parameters of LPMs. However, current PEFT methods encounter the challenge that GPU memory usage during training is not reduced as effectively as parameter usage. In this paper, we propose Fusion Side Tuning (FST), a novel memory-efficient and parameter-efficient fine-tuning method. FST significantly reduces GPU memory consumption during training, particularly at high resolutions. To achieve this, we freeze the backbone LPM and construct a learnable side fusion network that takes intermediate features from the backbone as input. The side fusion network consists of a sequence of fusion modules, which enable it to leverage the knowledge embedded in the intermediate features. Additionally, we employ an important token selection mechanism to further reduce training costs and memory requirements. We evaluate FST on eight medical image datasets of varying modalities and sizes. Experimental results demonstrate that FST outperforms existing PEFT methods, utilizing only 2% of the learnable parameters and 20% of the GPU memory required for full fine-tuning of a ViT-B encoder with an input resolution of 512 × 512.
Zhangchi Wang, Yijin Huang, Yidu Wu, Pujin Cheng, Li Lin 0006, Qinghai Guo, Xiaoying Tang 0001
BIBM6
2024 BKDSNN: Enhancing the Performance of Learning-Based Spiking Neural Networks Training with Blurred Knowledge Distillation
Kang You, Qinghai Guo, Zhezhi He
ECCV (50)3
2024 SpikeZIP-TF: Conversion is All You Need for Transformer-based SNN
abstract
Spiking neural network (SNN) has attracted great attention due to its characteristic of high efficiency and accuracy. Currently, the ANN-to-SNN conversion methods can obtain ANN on-par accuracy SNN with ultra-low latency (8 time-steps) in CNN structure on computer vision (CV) tasks. However, as Transformer-based networks have achieved prevailing precision on both CV and natural language processing (NLP), the Transformer-based SNNs are still encounting the lower accuracy w.r.t the ANN counterparts. In this work, we introduce a novel ANN-to-SNN conversion method called SpikeZIP-TF, where ANN and SNN are exactly equivalent, thus incurring no accuracy degradation. SpikeZIP-TF achieves 83.82% accuracy on CV dataset (ImageNet) and 93.79% accuracy on NLP dataset (SST-2), which are higher than SOTA Transformer-based SNNs. The code is available in GitHub: https://github.com/Intelligent-Computing-Research-Group/SpikeZIP_transformer
Kang You, Chen Nie, Zhijie Deng, Qinghai Guo, Zhezhi He
ICML5
2024 Sample as you Infer: Predictive Coding with Langevin Dynamics
abstract
We present Langevin Predictive Coding (LPC), a novel algorithm for deep generative model learning that builds upon the predictive coding framework of computational neuroscience. By injecting Gaussian noise into the predictive coding inference procedure and incorporating an encoder network initialization, we reframe the approach as an amortized Langevin sampling method for optimizing a tight variational lower bound. To increase robustness to sampling step size, we present a lightweight preconditioning technique inspired by Riemannian Langevin methods and adaptive SGD. We compare LPC against VAEs by training generative models on benchmark datasets; our experiments demonstrate superior sample quality and faster convergence for LPC in a fraction of SGD training iterations, while matching or exceeding VAE performance across key metrics like FID, diversity and coverage.
Umais Zahid, Qinghai Guo, Zafeirios Fountas
ICML2
2024 SpikingMiniLM: energy-efficient spiking transformer for natural language understanding
Jiangrong Shen, Zeke Wang, Qinghai Guo, Rui Yan 0005, Gang Pan 0001, Huajin Tang
Sci. China Inf. Sci.4
2024 Learning with sparse reward in a gap junction network inspired by the insect mushroom body
abstract
Animals can learn in real-life scenarios where rewards are often only available when a goal is achieved. This 'distal' or 'sparse' reward problem remains a challenge for conventional reinforcement learning algorithms. Here we investigate an algorithm for learning in such scenarios, inspired by the possibility that axo-axonal gap junction connections, observed in neural circuits with parallel fibres such as the insect mushroom body, could form a resistive network. In such a network, an active node represents the task state, connections between nodes represent state transitions and their connection to actions, and current flow to a target state can guide decision making. Building on evidence that gap junction weights are adaptive, we propose that experience of a task can modulate the connections to form a graph encoding the task structure. We demonstrate that the approach can be used for efficient reinforcement learning under sparse rewards, and discuss whether it is plausible as an account of the insect mushroom body.
Tianqi Wei 0001, Qinghai Guo, Barbara Webb
PLoS Comput. Biol.2
2023 Neuro-Modulated Hebbian Learning for Fully Test-Time Adaptation
abstract
Fully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradation problem of deep neural networks. We take inspiration from the biological plausibility learning where the neuron responses are tuned based on a local synapse-change procedure and activated by competitive lateral inhibition rules. Based on these feed-forward learning rules, we design a soft Hebbian learning process which provides an unsupervised and effective mechanism for online adaptation. We observe that the performance of this feed-forward Hebbian learning for fully test-time adaptation can be significantly improved by incorporating a feedback neuromodulation layer. It is able to fine-tune the neuron responses based on the external feedback generated by the error backpropagation from the top inference layers. This leads to our proposed neuro-modulated Hebbian learning (NHL) method for fully test-time adaptation. With the unsupervised feed-forward soft Hebbian learning being combined with a learned neuromodulator to capture feedback from external responses, the source model can be effectively adapted during the testing process. Experimental results on benchmark datasets demonstrate that our proposed method can significantly improve the adaptation performance of network models and outperforms existing state-of-the-art methods.
Yushun Tang, Ce Zhang 0009, Shuoshuo Chen, Luziwei Leng, Qinghai Guo, Zhihai He
CVPR7
2023 Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention Modeling
abstract
Weakly-supervised action localization aims to recognize and localize action instancese in untrimmed videos with only video-level labels. Most existing models rely on multiple instance learning(MIL), where the predictions of unlabeled instances are supervised by classifying labeled bags. The MIL-based methods are relatively well studied with cogent performance achieved on classification but not on localization. Generally, they locate temporal regions by the video-level classification but overlook the temporal variations of feature semantics. To address this problem, we propose a novel attention-based hierarchically-structured latent model to learn the temporal variations of feature semantics. Specifically, our model entails two components, the first is an unsupervised change-points detection module that detects change-points by learning the latent representations of video features in a temporal hierarchy based on their rates of change, and the second is an attention-based classification model that selects the change-points of the foreground as the boundaries. To evaluate the effectiveness of our model, we conduct extensive experiments on two benchmark datasets, THUMOS-14 and ActivityNet-v1.3. The experiments show that our method outperforms current state-of-the-art methods, and even achieves comparable performance with fully-supervised methods.
Guiqin Wang, Peng Zhao 0001, Cong Zhao 0001, Shusen Yang, Luziwei Leng, Jianxing Liao, Qinghai Guo
ICCV8
2023 Cross-Inferential Networks for Source-Free Unsupervised Domain Adaptation
abstract
One central challenge in source-free unsupervised domain adaptation (UDA) is the lack of an effective approach to evaluate the prediction results of the adapted network model in the target domain. To address this challenge, we propose to explore a new method called cross-inferential networks (CIN). Our main idea is that, when we adapt the network model to predict the sample labels from encoded features, we use these prediction results to construct new training samples with derived labels to learn a new examiner network that performs a different but compatible task in the target domain. Specifically, in this work, the base network model is performing image classification while the examiner network is tasked to perform relative ordering of triplets of samples whose training labels are carefully constructed from the prediction results of the base network model. Two similarity measures, cross-network correlation matrix similarity and attention consistency, are then developed to provide important guidance for the UDA process. Our experimental results on benchmark datasets demonstrate that our proposed CIN approach can significantly improve the performance of source-free UDA.
Yushun Tang, Qinghai Guo, Zhihai He
ICIP2
2023 Hebbian Deep Learning Without Feedback
Adrien Journé, Hector Garcia Rodriguez, Qinghai Guo, Timoleon Moraitis
ICLR3
2023 Learning and processing the ordinal information of temporal sequences in recurrent neural circuits
abstract
Temporal sequence processing is fundamental in brain cognitive functions. Experimental data has indicated that the representations of ordinal information and contents of temporal sequences are disentangled in the brain, but the neural mechanism underlying this disentanglement remains largely unclear. Here, we investigate how recurrent neural circuits learn to represent the abstract order structure of temporal sequences, and how this disentangled representation of order structure from that of contents facilitates the processing of temporal sequences. We show that with an appropriate learn protocol, a recurrent neural circuit can learn a set of tree-structured attractor states to encode the corresponding tree-structured orders of given temporal sequences. This abstract temporal order template can then be bound with different contents, allowing for flexible and robust temporal sequence processing. Using a transfer learning task, we demonstrate that the reuse of a temporal order template facilitates the acquisition of new temporal sequences of the same or similar ordinal structure. Using a key-word spotting task, we demonstrate that the attractor representation of order structure improves the robustness of temporal sequence discrimination, if the ordinal information is the key to differentiate different sequences. We hope this study gives us insights into the neural mechanism of representing the ordinal information of temporal sequences in the brain, and helps us to develop brain-inspired temporal sequence processing algorithms.
Xiaolong Zou, Zhikun Chu, Qinghai Guo, Bo Ho, Si Wu 0001, Yuanyuan Mi
NeurIPS3
2023 Predictive Coding as a Neuromorphic Alternative to Backpropagation: A Critical Evaluation
abstract
Backpropagation has rapidly become the workhorse credit assignment algorithm for modern deep learning methods. Recently, modified forms of predictive coding (PC), an algorithm with origins in computational neuroscience, have been shown to result in approximately or exactly equal parameter updates to those under backpropagation. Due to this connection, it has been suggested that PC can act as an alternative to backpropagation with desirable properties that may facilitate implementation in neuromorphic systems. Here, we explore these claims using the different contemporary PC variants proposed in the literature. We obtain time complexity bounds for these PC variants, which we show are lower bounded by backpropagation. We also present key properties of these variants that have implications for neurobiological plausibility and their interpretations, particularly from the perspective of standard PC as a variational Bayes algorithm for latent probabilistic models. Our findings shed new light on the connection between the two learning frameworks and suggest that in its current forms, PC may have more limited potential as a direct replacement of backpropagation than previously envisioned.
Umais Zahid, Qinghai Guo, Zafeirios Fountas
Neural Comput.2
2022 Replay-Oriented Gradient Projection Memory for Continual Learning in Medical Scenarios
abstract
Despite the tremendous progress recently achieved by deep learning (DL) in medical image analysis, most DL models only concentrate on single data distribution, which follows the independent and identically distributed (i.i.d) assumption. However, in practice, image data distribution changes with clinical conditions, such as different scanner manufacturers, imaging settings, and statistics regions. Although one can further train the model on new data samples, updating a model with data from an unknown distribution will always result in the model’s performance degradation on the learned data, a notorious phenomenon called catastrophic forgetting. Therefore affects the applicability of DL algorithms in continuously changing clinical scenarios. In this study, we have proposed a new method to address the impact of changing distributions in continual learning scenarios and alleviate catastrophic forgetting. A gradient regularization approach is used to suppress forgetting, and a replay-oriented consistency calculation method combined with a subspace weighting strategy is proposed to improve the model plasticity further. The proposed replay-oriented gradient projection memory (RO-GPM) is evaluated on multiple fundus disease diagnosis datasets including a real-world application and a continual learning benchmark. The quantitative and visualization results demonstrate that the proposed RO-GPM achieves superior performance to state-of-the-art algorithms by a large margin.1
Kuang Shu, Heng Li 0010, Qinghai Guo, Luziwei Leng, Jianxing Liao, Jiang Liu 0001
BIBM4
2022 Discrete time convolution for fast event-based stereo
abstract
Inspired by biological retina, dynamical vision sensor transmits events of instantaneous changes of pixel intensity, giving it a series of advantages over traditional frame-based camera, such as high dynamical range, high temporal resolution and low power consumption. However, extracting information from highly asynchronous event data is a challenging task. Inspired by continuous dynamics of biological neuron models, we propose a novel encoding method for sparse events-continuous time convolution (CTC)-which learns to model the spatial feature of the data with intrinsic dynamics. Adopting channel-wise parameterization, temporal dynamics of the model is synchronized on the same feature map and diverges across different ones, enabling it to embed data in a variety of temporal scales. Abstracted from CTC, we further develop discrete time convolution (DTC) which accelerates the process with lower computational cost. We apply these methods to event-based multi- view stereo matching where they surpass state-of-the-art methods on benchmark criteria of the MVSEC dataset. Spatially sparse event data often leads to inaccurate estimation of edges and local contours. To address this problem, we propose a dual-path architecture in which the feature map is complemented by underlying edge information from original events extracted with spatially-adaptive denormal-ization. We demonstrate the superiority of our model in terms of speed (up to 110 FPS), accuracy and robustness, showing a great potential for real-time fast depth estimation. Finally, we perform experiments on the recent DSEC dataset to demonstrate the general usage of our model.
Kaiwei Che, Jianguo Zhang 0001, Qinghai Guo, Luziwei Leng
CVPR6
2022 Spike-inspired rank coding for fast and accurate recurrent neural networks
Alan Jeffares, Qinghai Guo, Pontus Stenetorp, Timoleon Moraitis
ICLR2
2022 Variational Predictive Routing with Nested Subjective Timescales
Alexey Zakharov, Qinghai Guo, Zafeirios Fountas
ICLR2
2022 Short-Term Plasticity Neurons Learning to Learn and Forget
abstract
Short-term plasticity (STP) is a mechanism that stores decaying memories in synapses of the cerebral cortex. In computing practice, STP has been used, but mostly in the niche of spiking neurons, even though theory predicts that it is the optimal solution to certain dynamic tasks. Here we present a new type of recurrent neural unit, the STP Neuron (STPN), which indeed turns out strikingly powerful. Its key mechanism is that synapses have a state, propagated through time by a self-recurrent connection-within-the-synapse. This formulation enables training the plasticity with backpropagation through time, resulting in a form of learning to learn and forget in the short term. The STPN outperforms all tested alternatives, i.e. RNNs, LSTMs, other models with fast weights, and differentiable plasticity. We confirm this in both supervised and reinforcement learning (RL), and in tasks such as Associative Retrieval, Maze Exploration, Atari video games, and MuJoCo robotics. Moreover, we calculate that, in neuromorphic or biological circuits, the STPN minimizes energy consumption across models, as it depresses individual synapses dynamically. Based on these, biological STP may have been a strong evolutionary attractor that maximizes both efficiency and computational power. The STPN now brings these neuromorphic advantages also to a broad spectrum of machine learning practice. Code is available in https://github.com/NeuromorphicComputing/stpn.
Hector Garcia Rodriguez, Qinghai Guo, Timoleon Moraitis
ICML2
2022 Differentiable hierarchical and surrogate gradient search for spiking neural networks
abstract
Spiking neural network (SNN) has been viewed as a potential candidate for the next generation of artificial intelligence with appealing characteristics such as sparse computation and inherent temporal dynamics. By adopting architectures of deep artificial neural networks (ANNs), SNNs are achieving competitive performances in benchmark tasks such as image classification. However, successful architectures of ANNs are not necessary ideal for SNN and when tasks become more diverse effective architectural variations could be critical. To this end, we develop a spike-based differentiable hierarchical search (SpikeDHS) framework, where spike-based computation is realized on both the cell and the layer level search space. Based on this framework, we find effective SNN architectures under limited computation cost. During the training of SNN, a suboptimal surrogate gradient function could lead to poor approximations of true gradients, making the network enter certain local minima. To address this problem, we extend the differential approach to surrogate gradient search where the SG function is efficiently optimized locally. Our models achieve state-of-the-art performances on classification of CIFAR10/100 and ImageNet with accuracy of 95.50%, 76.25% and 68.64%. On event-based deep stereo, our method finds optimal layer variation and surpasses the accuracy of specially designed ANNs meanwhile with 26$\times$ lower energy cost ($6.7\mathrm{mJ}$), demonstrating the advantage of SNN in processing highly sparse and dynamic signals. Codes are available at \url{https://github.com/Huawei-BIC/SpikeDHS}.
Kaiwei Che, Luziwei Leng, Jianguo Zhang 0001, Qinghu Meng, Qinghai Guo, Jianxing Liao
NeurIPS7
2022 Self-Supervised Learning Through Efference Copies
abstract
Self-supervised learning (SSL) methods aim to exploit the abundance of unlabelled data for machine learning (ML), however the underlying principles are often method-specific. An SSL framework derived from biological first principles of embodied learning could unify the various SSL methods, help elucidate learning in the brain, and possibly improve ML. SSL commonly transforms each training datapoint into a pair of views, uses the knowledge of this pairing as a positive (i.e. non-contrastive) self-supervisory sign, and potentially opposes it to unrelated, (i.e. contrastive) negative examples. Here, we show that this type of self-supervision is an incomplete implementation of a concept from neuroscience, the Efference Copy (EC). Specifically, the brain also transforms the environment through efference, i.e. motor commands, however it sends to itself an EC of the full commands, i.e. more than a mere SSL sign. In addition, its action representations are likely egocentric. From such a principled foundation we formally recover and extend SSL methods such as SimCLR, BYOL, and ReLIC under a common theoretical framework, i.e. Self-supervision Through Efference Copies (S-TEC). Empirically, S-TEC restructures meaningfully the within- and between-class representations. This manifests as improvement in recent strong SSL baselines in image classification, segmentation, object detection, and in audio. These results hypothesize a testable positive influence from the brain's motor outputs onto its sensory representations.
Franz Scherr, Qinghai Guo, Timoleon Moraitis
NeurIPS2
2022 DevFly: Bio-Inspired Development of Binary Connections for Locality Preserving Sparse Codes
abstract
Neural circuits undergo developmental processes which can be influenced by experience. Here we explore a bio-inspired development process to form the connections in a network used for locality sensitive hashing. The network is a simplified model of the insect mushroom body, which has sparse connections from the input layer to a second layer of higher dimension, forming a sparse code. In previous versions of this model, connectivity between the layers is random. We investigate whether the performance of the hash, evaluated in nearest neighbour query tasks, can be improved by process of developing the connections, in which the strongest input dimensions in successive samples are wired to each successive coding dimension. Experiments show that the accuracy of searching for nearest neighbours is improved, although performance is dependent on the parameter values and datasets used. Our approach is also much faster than alternative methods that have been proposed for training the connections in this model. Importantly, the development process does not impact connections built at an earlier stage, which should provide stable coding results for simultaneous learning in a downstream network.
Tianqi Wei 0001, Rana Alkhoury Maroun, Qinghai Guo, Barbara Webb
NeurIPS3