VLDB 2026 Research / reviewers in the wild / expert
Guoqi Li 0002
dblp:13/5353-2
· DBLP profile ↗
102ranked-venue papers
10as first author
71since 2021 · last 2026
0000-0002-8994-431XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 7 first-author · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language ModelsabstractXingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li, Jiajun Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xingrun Xing, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li 0002 |
ACL (1) | 8 |
| 2026 | Event-based low-power spiking gaze estimation
Zhipeng Sui, Weihua He, Yongxiang Feng, Xiaobao Wei, Qiushuang Lian, Guoqi Li 0002, Wenhui Wang 0001 |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | FastGaze: An efficient and flexible model for human scanpath prediction
Jiahong Zhang, Hongjuan Pei, Tianxiang Hu, Richard D. Shang, Bo Xu 0002, Guoqi Li 0002 |
Expert Syst. Appl. | 8 |
| 2026 | Advancing the forward-forward algorithm towards high-performance deep local learning
Yujie Wu 0002, Jibin Wu, Lei Deng 0003, Mingkun Xu, Qinghao Wen, Guoqi Li 0002 |
Neural Networks | 7 |
| 2026 | Enhancing robustness of spiking neural networks through retina-like coding and memory-based neurons
Jiahong Zhang, Man Yao, Peng Zhou 0017, Bo Xu 0002, Guoqi Li 0002 |
Neural Networks | 7 |
| 2025 | Spike2Former: Efficient Spiking Transformer for High-performance Image SegmentationabstractSpiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0x efficiency on ADE20K, +14.3% mIoU and 5.2x efficiency on VOC2012, and +9.1% mIoU and 6.6x efficiency on CityScapes. Zhenxin Lei, Man Yao, Xinhao Luo, Yanye Lu, Bo Xu 0002, Guoqi Li 0002 |
AAAI | 7 |
| 2025 | Efficient 3D Recognition with Event-driven Spike Sparse ConvolutionabstractSpiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature. Xuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou, Shibo Zhou, Bo Xu 0002, Guoqi Li 0002 |
AAAI | 8 |
| 2025 | MMDEND: Dendrite-Inspired Multi-Branch Multi-Compartment Parallel Spiking Neuron for Sequence ModelingabstractKexin Wang, Yuhong Chou, Di Shang, Shijie Mei, Jiahong Zhang, Yanbin Huang, Man Yao, Bo Xu, Guoqi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuhong Chou, Richard D. Shang, Shijie Mei 0001, Jiahong Zhang, Yanbin Huang, Man Yao, Bo Xu 0002, Guoqi Li 0002 |
ACL (1) | 9 |
| 2025 | MVA: Linear Attention with High-order Query-Keys Integration and Multi-level Vocabulary DecompositionabstractLinear attention offers the advantages of linear inference time and fixed memory usage compared to Softmax attention.
However, training large-scale language models with linear attention from scratch remains prohibitively expensive and exhibits significant performance gaps compared to Softmax-based models.
To address these challenges, we focus on transforming pre-trained Softmax-based language models into linear attention models.
We unify mainstream linear attention methods using a **high-order QK integration theory** and a **multi-level vocabulary decomposition**.
Specifically, the QK integration theory explains the efficacy of combining linear and sparse attention from the perspective of information collection across different frequency bands.
The multi-level vocabulary decomposition exponentially expands memory capacity by recursively exploiting compression loss from compressed states.
Through detailed error analysis, we demonstrate superior approximation of Softmax attention achieved by our approach.
To further improve performance and reduce training costs, we adopt a **soft integration strategy** with attention scores, effectively combining a sliding window mechanism.
With less than 100M tokens, our method fine-tunes models to achieve linear complexity while retaining 99\% of their original performance.
Compared to state-of-the-art linear attention model and method, our approach improves MMLU scores by 1.2 percentage points with minimal fine-tuning.
Furthermore, even without the sliding window mechanism, our method achieves state-of-the-art performance on all test sets with 10B tokens. Zekun Li 0014, Tongxin Bai, Man Yao, Guoqi Li 0002 |
ICML | 6 |
| 2025 | Enabling scale and rotation invariance in convolutional neural networks with retina like transformation
Jiahong Zhang, Guoqi Li 0002, Qiaoyi Su, Lihong Cao, Yonghong Tian 0001, Bo Xu 0002 |
Neural Networks | 2 |
| 2025 | Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation TrainingabstractThe ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5%, 79.8%, 84.0%, and 86.2% with models containing 10 M, 19 M, 83 M, and 173 M parameters, respectively. For instance, the 10 M model outperforms the best existing SNN by 7.2% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5× and 3.9×, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone. Man Yao, Xuerui Qiu, Tianxiang Hu, Yuhong Chou, Keyu Tian, Jianxing Liao, Luziwei Leng, Bo Xu 0002, Guoqi Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | Event-Based Video Reconstruction Via Spatial-Temporal Heterogeneous Spiking Neural NetworkabstractEvent cameras detect per-pixel brightness changes and output asynchronous event streams with high temporal resolution, high dynamic range, and low latency. However, the unstructured nature of event streams means that humans cannot analyze and interpret them in the same way as natural images. Event-based video reconstruction is a widely used method aimed at reconstructing intuitive videos from event streams. Most reconstruction methods based on traditional artificial neural networks (ANNs) have high energy consumption, which counteracts the low-power advantage of event cameras. Spiking neural networks (SNNs) are a new generation of event-driven neural networks that encode information via discrete spikes, which leads to greater computational efficiency. Previous methods based on SNNs overlooked the asynchronous nature of event streams, leading to reconstructions that suffer from artifacts, flickering, low contrast, etc. In this work, we analyze event streams and spiking neurons and explain poor reconstruction quality. We specifically propose a novel spatial-temporal heterogeneous (STH) spiking neuron suitable for reconstructing asynchronous event streams. The STH neuron adjusts the membrane decay coefficient adaptively and has better spatiotemporal perception. In addition, we propose a temporal-frequency calibration module (TFCM) based on the Fourier transform to improve the contrast of the reconstructions. On the basis of the above proposed neuron and module, we construct two SNN-based models, referred to as the STHSNN and TFCSNN. The goal of the former is to reduce the artifacts and flickering in reconstructions, whereas the latter focuses on enhancing the contrast. The experimental results demonstrate that our models can yield reconstructions in various scenarios, achieving better quality and lower energy consumption than previous SNNs. Specifically, the TFCSNN and STHSNN achieve top-2 performance among the SNN-based models, with energy consumption reductions of 3.48 times and 12.40 times, respectively. Lijun Guo, Chong Wang 0001, Guoqi Li 0002, Jiangbo Qian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | DAFDet: A Unified Dynamic SAR Target Detection Architecture With Asymptotic Fusion Enhancement and Feature Encoding DecouplingabstractIn many military and civilian applications, synthetic aperture radar (SAR) image target detection plays a vital role. However, current methods for SAR target detection generally fail to balance speed and accuracy, thus making it impossible to deploy them to real-world engineering applications. In addition, strong scattering, multiscales, high density, complex background interference, and speckle noise make it remarkably challenging to extract effective target information and disentangle background noise from the target information, ultimately resulting in high missing and false alarm rates. To address these issues, a unified dynamic SAR target detection architecture (DAFDet) with asymptotic fusion enhancement and feature encoding decoupling is proposed in this article. First, a dynamic architecture is constructed by cascading two identical detectors and integrating a designed decision maker. This decision maker can automatically decide the inference route by calculating the difficulty score of an SAR image, which ensures efficient inference speed while achieving high accuracy. Second, an asymptotic fusion enhancement feature pyramid network (AFEFPN) is developed, which can avoid the loss and degradation of target information in multistage transmissions through direct interactions of nonadjacent levels, thereby enhancing the extraction of valid target information. Suppression of background noise is achieved by modeling the importance of different feature channels of the fused features. Finally, a task-oriented decoupled head (TODH) is proposed to boost the localization and classification abilities of the model in complex scenarios. It decouples feature encoding at the source, thus providing task-oriented feature context. Numerous experiments on four widely adopted datasets reveal that DAFDet obtains efficient inference speed and optimal detection accuracy, achieving new state-of-the-art target detection performance. The source code will be provided athttps://github.com/yangyahu-1994/DAFDet. Yahu Yang, Yuntao Du 0006, Li Zhang 0025, Guoqi Li 0002, Yushi Chen 0002, Guorui Cheng, Shenmin Song |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | An Efficient Sequential Decentralized Federated Progressive Channel Pruning Strategy for Smart Grid Electricity Theft DetectionabstractThis article aims to develop a lightweight, decentralized federated learning (FL)-based strategy for electricity theft detection (ETD). Different from most of the existing ETD solutions, which typically deploy centralized deep learning models, our proposed method utilizes well-pruned lightweight networks and operates in a completely decentralized manner while maintaining the performance of the ETD model. Specifically, to protect data privacy, a novel sequential decentralized FL (SDFL) framework was designed, eliminating the centralized parameter aggregation node in traditional FL. Each client communicates model parameters only with its neighbors and trains its model locally. In addition, to facilitate deployment on edge devices, model pruning techniques are integrated with the sequential transmission characteristics of the SDFL framework. A progressive channel pruning technique is proposed, gradually reducing the number of model channels during training to promote model compression and simplify field deployment. Experiments demonstrate that our strategy compressed the model floating point operations from 18.32 to 3.60M and reduced the number of parameters from 8.61 to 3.47M, while protecting user privacy, and maintaining good performance. Deployment results on the edge devices, i.e., Raspberry Pi, indicate that our proposed strategy reduces the model inference time from 329.35 to 141.50 s, enhancing the detection efficiency by 57.04%. Fanghong Guo, Hao Yang 0047, Guoqi Li 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | AuthSim: Toward Authentic and Effective Safety-Critical Scenario Generation for Autonomous Driving TestsabstractThe generation of adversarial safety-critical scenarios is essential for rigorously evaluating autonomous driving systems, enabling the identification of vulnerabilities and enhancement of system robustness. However, existing methodologies predominantly focus on extreme, unconstrained collision scenarios in which non-player character (NPC) vehicles exhibit unrealistic adversarial behaviors toward the ego vehicle. While such scenarios serve as stress tests, their practical utility is limited due to two key factors: 1) these extreme events are statistically rare in real-world traffic and frequently involve collisions that are physically unavoidable, irrespective of the autonomous vehicle’s decision-making capabilities; and 2) NPC behaviors in these scenarios are often intentionally aggressive (e.g., deliberate rear-end collisions), resulting in liability attribution that predominantly lies with the NPCs rather than exposing meaningful system limitations. Recent efforts to enhance scenario plausibility rely extensively on large-scale real-world traffic datasets, introducing significant computational costs and scalability constraints. To overcome these limitations, we propose a three-layer relative safety region model that partitions the driving environment into zones of varying risk levels. This partitioning increases the likelihood that NPC vehicles will interact within relative safety boundary regions, thus enabling the generation of more realistic and contextually relevant adversarial scenarios without the need for extensive real-world traffic data. We introduce AuthSim, a platform that integrates this safety model with reinforcement learning (RL) to generate both authentic and effective safety-critical scenarios. AuthSim is the first comprehensive approach to address both the authenticity and effectiveness of autonomous driving test scenarios without relying on large-scale traffic data. Empirical results demonstrate that AuthSim outperforms existing methods, achieving a 5.25% improvement in average cut-in distance and a 11.94% increase in average collision interval time compared with the state-of-the-art (SOTA) results, all while maintaining superior efficiency in scenario generation. These findings highlight the potential of AuthSim to produce high-fidelity and efficient test cases for the rigorous evaluation of autonomous driving systems. Yukuan Yang, Xucheng Lu, Zepeng Wu, Guoqi Li 0002, Lingzhong Meng, Zhiming Ding, Yunzhi Xue |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Advancing Spiking Neural Networks Toward Deep Residual LearningabstractDespite the rapid progress of neuromorphic computing, inadequate capacity and insufficient representation power of spiking neural networks (SNNs) severely restrict their application scope in practice. Residual learning and shortcuts have been evidenced as an important approach for training deep neural networks, but rarely did previous work assessed their applicability to the specifics of SNNs. In this article, we first identify that this negligence leads to impeded information flow and the accompanying degradation problem in a spiking version of vanilla ResNet. To address this issue, we propose a novel SNN-oriented residual architecture termed MS-ResNet, which establishes membrane-based shortcut pathways, and further proves that the gradient norm equality can be achieved in MS-ResNet by introducing block dynamical isometry theory, which ensures the network can be well-behaved in a depth-insensitive way. Thus, we are able to significantly extend the depth of directly trained SNNs, e.g., up to 482 layers on CIFAR-10 and 104 layers on ImageNet, without observing any slight degradation problem. To validate the effectiveness of MS-ResNet, experiments on both frame-based and neuromorphic datasets are conducted. MS-ResNet104 achieves a superior result of 76.02% accuracy on ImageNet, which is the highest to the best of our knowledge in the domain of directly trained SNNs. Great energy efficiency is also observed, with an average of only one spike per neuron needed to classify an input sample. We believe our powerful and scalable models will provide strong support for further exploration of SNNs. Yifan Hu 0013, Lei Deng 0003, Yujie Wu 0002, Man Yao, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Gated Attention Coding for Training High-Performance and Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are emerging as an energy-efficient alternative to traditional artificial neural networks (ANNs) due to their unique spike-based event-driven nature. Coding is crucial in SNNs as it converts external input stimuli into spatio-temporal feature sequences. However, most existing deep SNNs rely on direct coding that generates powerless spike representation and lacks the temporal dynamics inherent in human vision. Hence, we introduce Gated Attention Coding (GAC), a plug-and-play module that leverages the multi-dimensional gated attention unit to efficiently encode inputs into powerful representations before feeding them into the SNN architecture. GAC functions as a preprocessing layer that does not disrupt the spike-driven nature of the SNN, making it amenable to efficient neuromorphic hardware implementation with minimal modifications. Through an observer model theoretical analysis, we demonstrate GAC's attention mechanism improves temporal dynamics and coding efficiency. Experiments on CIFAR10/100 and ImageNet datasets demonstrate that GAC achieves state-of-the-art accuracy with remarkable efficiency. Notably, we improve top-1 accuracy by 3.10% on CIFAR100 with only 6-time steps and 1.07% on ImageNet while reducing energy usage to 66.9% of the previous works. To our best knowledge, it is the first time to explore the attention-based dynamic coding scheme in deep SNNs, with exceptional effectiveness and efficiency on large-scale datasets. Code is available at https://github.com/bollossom/GAC. Xuerui Qiu, Rui-Jie Zhu 0003, Yuhong Chou, Zhaorui Wang 0005, Liang-Jian Deng, Guoqi Li 0002 |
AAAI | 6 |
| 2024 | HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsabstractThe main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS. Xiao Wang 0014, Zongzhen Wu, Bo Jiang 0002, Zhimin Bao, Lin Zhu 0012, Guoqi Li 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
AAAI | 6 |
| 2024 | SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural NetworkabstractBrain-inspired Spiking Neural Network (SNN) has demonstrated its effectiveness and efficiency in vision, natural language, and speech understanding tasks, indicating their capacity to “see”, “listen”, and “read”. In this paper, we design SpikeVoice, which performs high-quality Text-To-Speech (TTS) via SNN, to explore the potential of SNN to “speak”. A major obstacle to using SNN for such generative tasks lies in the demand for models to grasp long-term dependencies. The serial nature of spiking neurons, however, leads to the invisibility of information at future spiking time steps, limiting SNN models to capture sequence dependencies solely within the same time step. We term this phenomenon “partial-time dependency”. To address this issue, we introduce Spiking Temporal-Sequential Attention (STSA) in the SpikeVoice. To the best of our knowledge, SpikeVoice is the first TTS work in the SNN field. We perform experiments using four well-established datasets that cover both Chinese and English languages, encompassing scenarios with both single-speaker and multi-speaker configurations. The results demonstrate that SpikeVoice can achieve results comparable to Artificial Neural Networks (ANN) with only 10.5% energy consumption of ANN. Both our demo and code are available as supplementary material. Jiahong Zhang, Yong Ren 0006, Man Yao, Richard D. Shang, Bo Xu 0002, Guoqi Li 0002 |
ACL (1) | 7 |
| 2024 | Integer-Valued Training and Spike-Driven Inference Spiking Neural Network for High-Performance and Energy-Efficient Object Detection
Xinhao Luo, Man Yao, Yuhong Chou, Bo Xu 0002, Guoqi Li 0002 |
ECCV (32) | 5 |
| 2024 | Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic ChipsabstractNeuromorphic computing, which exploits Spiking Neural Networks (SNNs) on neuromorphic chips, is a promising energy-efficient alternative to traditional AI. CNN-based SNNs are the current mainstream of neuromorphic computing. By contrast, no neuromorphic chips are designed especially for Transformer-based SNNs, which have just emerged, and their performance is only on par with CNN-based SNNs, offering no distinct advantage. In this work, we propose a general Transformer-based SNN architecture, termed as ``Meta-SpikeFormer", whose goals are: (1) *Lower-power*, supports the spike-driven paradigm that there is only sparse addition in the network; (2) *Versatility*, handles various vision tasks; (3) *High-performance*, shows overwhelming performance advantages over CNN-based SNNs; (4) *Meta-architecture*, provides inspiration for future next-generation Transformer-based neuromorphic chip designs. Specifically, we extend the Spike-driven Transformer in \citet{yao2023spike} into a meta architecture, and explore the impact of structure, spike-driven self-attention, and skip connection on its performance. On ImageNet-1K, Meta-SpikeFormer achieves 80.0\% top-1 accuracy (55M), surpassing the current state-of-the-art (SOTA) SNN baselines (66M) by 3.7\%. This is the first direct training SNN backbone that can simultaneously supports classification, detection, and segmentation, obtaining SOTA results in SNNs. Finally, we discuss the inspiration of the meta SNN architecture for neuromorphic chip design. Man Yao, Tianxiang Hu, Zhaokun Zhou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
ICLR | 8 |
| 2024 | High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost
Man Yao, Xuerui Qiu, Yuhong Chou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
ICML | 9 |
| 2024 | RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
Keming Wu, Man Yao, Yuhong Chou, Xuerui Qiu, Bo Xu 0002, Guoqi Li 0002 |
ACM Multimedia | 7 |
| 2024 | MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapabstractVarious linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: (1) Dynamic memory ability; (2) Static approximation ability; (3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models. Yuhong Chou, Man Yao, Yuqi Pan, Rui-Jie Zhu 0003, Jibin Wu, Yiran Zhong, Bo Xu 0002, Guoqi Li 0002 |
NeurIPS | 10 |
| 2024 | Spiking Transformer with Experts MixtureabstractSpiking Neural Networks (SNNs) provide a sparse spike-driven mechanism which is believed to be critical for energy-efficient deep learning.
Mixture-of-Experts (MoE), on the other side, aligns with the brain mechanism of distributed and sparse processing, resulting in an efficient way of enhancing model capacity and conditional computation.
In this work, we consider how to incorporate SNNs’ spike-driven and MoE’s conditional computation into a unified framework.
However, MoE uses softmax to get the dense conditional weights for each expert and TopK to hard-sparsify the network, which does not fit the properties of SNNs.
To address this issue, we reformulate MoE in SNNs and introduce the Spiking Experts Mixture Mechanism (SEMM) from the perspective of sparse spiking activation.
Both the experts and the router output spiking sequences, and their element-wise operation makes SEMM computation spike-driven and dynamic sparse-conditional.
By developing SEMM into Spiking Transformer, the Experts Mixture Spiking Attention (EMSA) and the Experts Mixture Spiking Perceptron (EMSP) are proposed, which performs routing allocation for head-wise and channel-wise spiking experts, respectively. Experiments show that SEMM realizes sparse conditional computation and obtains a stable improvement on neuromorphic and static datasets with approximate computational overhead based on the Spiking Transformer baselines. Zhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che, Liwei Huang, Yuesheng Zhu, Guoqi Li 0002, Zhaofei Yu, Li Yuan 0007 |
NeurIPS | 9 |
| 2024 | Multi-scale full spike pattern for semantic segmentation
Qiaoyi Su, Weihua He, Xiaobao Wei, Bo Xu 0002, Guoqi Li 0002 |
Neural Networks | 5 |
| 2024 | SNN-BERT: Training-efficient Spiking Neural Networks for energy-efficient BERT
Qiaoyi Su, Shijie Mei 0001, Xingrun Xing, Man Yao, Bo Xu 0002, Guoqi Li 0002 |
Neural Networks | 7 |
| 2024 | Brain-Inspired Computing: A Systematic Survey and Future TrendsabstractBrain-inspired computing (BIC) is an emerging research field that aims to build fundamental theories, models, hardware architectures, and application systems toward more general artificial intelligence (AI) by learning from the information processing mechanisms or structures/functions of biological nervous systems. It is regarded as one of the most promising research directions for future intelligent computing in the post-Moore era. In the past few years, various new schemes in this field have sprung up to explore more general AI. These works are quite divergent in the aspects of modeling/algorithm, software tool, hardware platform, and benchmark data since BIC is an interdisciplinary field that consists of many different domains, including computational neuroscience, AI, computer science, statistical physics, material science, and microelectronics. This situation greatly impedes researchers from obtaining a clear picture and getting started in the right way. Hence, there is an urgent requirement to do a comprehensive survey in this field to help correctly recognize and analyze such bewildering methodologies. What are the key issues to enhance the development of BIC? What roles do the current mainstream technologies play in the general framework of BIC? Which techniques are truly useful in real-world applications? These questions largely remain open. To address the above issues, in this survey, we first clarify the biggest challenge of BIC: how can AI models benefit from the recent advancements in computational neuroscience? With this challenge in mind, we will focus on discussing the concept of BIC and summarize four components of BIC infrastructure development: 1) modeling/algorithm; 2) hardware platform; 3) software tool; and 4) benchmark data. For each component, we will summarize its recent progress, main challenges to resolve, and future trends. Based on these studies, we present a general framework for the real-world applications of BIC systems, which is promising to benefit both AI and brain science. Finally, we claim that it is extremely important to build a research ecology to promote prosperity continuously in this field. Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 1 |
| 2024 | Corrections to "Brain-Inspired Computing: A Systematic Survey and Future Trends"abstractPresents corrections to the paper, (Corrections to “Brain-Inspired Computing: A Systematic Survey and Future Trends”). Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 1 |
| 2024 | Optimal Control of Temporal Networks With Variable Input and Node-Source ConnectionabstractMany networked systems built upon real-life physical or social interactions have time-varying connections among individual units, where the temporal changes in connectivity and/or interaction strength lead to complicated dynamics. The temporal network model was proposed in the form of controlled linear dynamical systems acting in an ordered sequence of time intervals. One of the core challenges in network science is the control of networks and the optimization of the control strategy. However, most canonical frameworks for solving optimal control problems were established for static networks featuring constant topology. New theories and techniques are yet to be developed for the temporal networks, with an important case being that the input and the source-node connection are both variables. In this work, by formulating a quadratic energy cost without solving the Riccati differential equation, we show that the control effort can be reduced substantially by improving either the system trajectories or the input matrices. The two approaches are further combined in a coordinate descent framework, integrating linearly constrained quadratic programming, and a projected gradient descent method. Taken together, the results underline the potential of temporal networks as energy-efficient control systems and present strategies to improve the control input. Moreover, the proposed algorithms can serve as a starting point for future engineering of real-world temporal networks. Yukun Hao, Jiangshuai Huang, Changyun Wen, Guoqi Li 0002 |
IEEE Trans. Cybern. | 5 |
| 2024 | Boosting Zero-Shot Learning via Contrastive Optimization of Attribute RepresentationsabstractZero-shot learning (ZSL) aims to recognize classes that do not have samples in the training set. One representative solution is to directly learn an embedding function associating visual features with corresponding class semantics for recognizing new classes. Many methods extend upon this solution, and recent ones are especially keen on extracting rich features from images, e.g., attribute features. These attribute features are normally extracted within each individual image; however, the common traits for features across images yet belonging to the same attribute are not emphasized. In this article, we propose a new framework to boost ZSL by explicitly learning attribute prototypes beyond images and contrastively optimizing them with attribute-level features within images. Besides the novel architecture, two elements are highlighted for attribute representations: a new prototype generation module (PM) is designed to generate attribute prototypes from attribute semantics; a hard-example-based contrastive optimization scheme is introduced to reinforce attribute-level features in the embedding space. We explore two alternative backbones, CNN-based and transformer-based, to build our framework and conduct experiments on three standard benchmarks, Caltech-UCSD Birds-200-2011 (CUB), SUN attribute database (SUN), and animals with attributes 2 (AwA2). Results on these benchmarks demonstrate that our method improves the state of the art by a considerable margin. Our codes will be available at https://github.com/dyabel/CoAR-ZSL.git. Yu Du 0010, Miaojing Shi, Fangyun Wei, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Rethinking Pretraining as a Bridge From ANNs to SNNsabstractSpiking neural networks (SNNs) are known as typical kinds of brain-inspired models with their unique features of rich neuronal dynamics, diverse coding schemes, and low power consumption properties. How to obtain a high-accuracy model has always been the main challenge in the field of SNN. Currently, there are two mainstream methods, i.e., obtaining a converted SNN through converting a well-trained artificial NN (ANN) to its SNN counterpart or training an SNN directly. However, the inference time of a converted SNN is too long, while SNN training is generally very costly and inefficient. In this work, a new SNN training paradigm is proposed by combining the concepts of the two different training methods with the help of the pretrain technique and BP-based deep SNN training mechanism. We believe that the proposed paradigm is a more efficient pipeline for training SNNs. The pipeline includes pipe-S for static data transfer tasks and pipe-D for dynamic data transfer tasks. State-of-the-art (SOTA) results are obtained in a large-scale event-driven dataset ES-ImageNet. For training acceleration, we achieve the same (or higher) best accuracy as similar leaky-integrate-and-fire (LIF)-SNNs using 1/8 training time on ImageNet-1K and 1/2 training time on ES-ImageNet and also provide a time-accuracy benchmark for a new dataset ES-UCF101. These experimental results reveal the similarity of the functions of parameters between ANNs and SNNs and also demonstrate various potential applications of this SNN training pipeline. Yifan Hu 0013, Shijie Ma, Dongjie Yu, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Spike Attention Coding for Spiking Neural NetworksabstractSpiking neural networks (SNNs), an important family of neuroscience-oriented intelligent models, play an essential role in the neuromorphic computing community. Spike rate coding and temporal coding are the mainstream coding schemes in the current modeling of SNNs. However, rate coding usually suffers from limited representation resolution and long latency, while temporal coding usually suffers from under-utilization of spike activities. To this end, we propose spike attention coding (SAC) for SNNs. By introducing learnable attention coefficients for each time step, our coding scheme can naturally unify rate coding and temporal coding, and then flexibly learn optimal coefficients for better performance. Several normalization and regularization techniques are further incorporated to control the range and distribution of the learned attention coefficients. Extensive experiments on classification, generation, and regression tasks are conducted and demonstrate the superiority of the proposed coding scheme. This work provides a flexible coding scheme to enhance the representation power of SNNs and extends their application scope beyond the mainstream classification scenario. Yifan Hu 0013, Guoqi Li 0002, Jing Pei, Lei Deng 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Task-Prompt Generalised World Model in Multi-Environment Offline Reinforcement LearningabstractOffline reinforcement learning (RL) circumvents costly interactions with the environment by utilising historical trajectories. Incorporating a world model into this method could substantially enhance the transfer performance of various tasks without expensive calculations from scratch. However, due to the complexity arising from different types of generalisation, previous works have focused almost exclusively on single-environment tasks. In this study, we introduce a multi-environment offline RL setting to investigate whether a generalised world model can be learned from large, diverse datasets and serve as a good surrogate for policy learning in different tasks. Inspired by the success of multi-task prompt methods, we propose the Task-prompt Generalised World Model (TGW) framework, which demonstrates notable performance in this setting. TGW comprises three modules: a task-state prompter, a generalised dynamics module, and a reward module. We implement the generalised dynamics module as a transformer-based recurrent state-space model and employ prompts to provide task-specific instructions, enabling TGW to address the internal stochasticity of the generalised world model. On the MuJoCo control benchmarks, TGW significantly outperforms previous offline RL algorithms in multi-environment setting. Xuantang Xiong, Linghui Meng 0001, Jingqing Ruan, Qingyang Zhang 0004, Guoqi Li 0002, Dengpeng Xing, Bo Xu 0002 |
ECAI | 5 |
| 2023 | Deep Directly-Trained Spiking Neural Networks for Object DetectionabstractSpiking neural networks (SNNs) are brain-inspired energy-efficient models that encode information in spatiotemporal dynamics. Recently, deep SNNs trained directly have shown great success in achieving high performance on classification tasks with very few time steps. However, how to design a directly-trained SNN for the regression task of object detection still remains a challenging problem. To address this problem, we propose EMS-YOLO, a novel directly-trained SNN framework for object detection, which is the first trial to train a deep SNN with surrogate gradients for object detection rather than ANN-SNN conversion strategies. Specifically, we design a full-spike residual block, EMS-ResNet, which can effectively extend the depth of the directly-trained SNN with low power consumption. Furthermore, we theoretically analyze and prove the EMS-ResNet could avoid gradient vanishing or exploding. The results demonstrate that our approach outperforms the state-of-the-art ANN-SNN conversion methods (at least 500 time steps) in extremely fewer time steps (only 4 time steps). It is shown that our model could achieve comparable performance to the ANN with the same architecture while consuming 5.83× less energy on the frame-based COCO Dataset and the event-based Gen1 Dataset. Our code is available in https://github.com/BICLab/EMS-YOLO. Qiaoyi Su, Yuhong Chou, Yifan Hu 0013, Jianing Li 0001, Shijie Mei 0001, Guoqi Li 0002 |
ICCV | 7 |
| 2023 | Inherent Redundancy in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are well known as a promising energy-efficient alternative to conventional artificial neural networks. Subject to the preconceived impression that SNNs are sparse firing, the analysis and optimization of inherent redundancy in SNNs have been largely overlooked, thus the potential advantages of spike-based neuromorphic computing in accuracy and energy efficiency are interfered. In this work, we pose and focus on three key questions regarding the inherent redundancy in SNNs. We argue that the redundancy is induced by the spatio-temporal invariance of SNNs, which enhances the efficiency of parameter utilization but also invites lots of noise spikes. Further, we analyze the effect of spatio-temporal invariance on the spatio-temporal dynamics and spike firing of SNNs. Then, motivated by these analyses, we propose an Advance Spatial Attention (ASA) module to harness SNNs’ redundancy, which can adaptively optimize their membrane potential distribution by a pair of individual spatial attention sub-modules. In this way, noise spike features are accurately regulated. Experimental results demonstrate that the proposed method can significantly drop the spike firing with better performance than state-of-the-art SNN baselines. Our code is available in https://github.com/BICLab/ASA-SNN. Man Yao, Guang-She Zhao, Yaoyuan Wang, Bo Xu 0002, Guoqi Li 0002 |
ICCV | 7 |
| 2023 | Spike-driven TransformerabstractSpiking Neural Networks (SNNs) provide an energy-efficient deep learning option due to their unique spike-based event-driven (i.e., spike-driven) paradigm. In this paper, we incorporate the spike-driven paradigm into Transformer by the proposed Spike-driven Transformer with four unique properties: (1) Event-driven, no calculation is triggered when the input of Transformer is zero; (2) Binary spike communication, all matrix multiplications associated with the spike matrix can be transformed into sparse additions; (3) Self-attention with linear complexity at both token and channel dimensions; (4) The operations between spike-form Query, Key, and Value are mask and addition. Together, there are only sparse addition operations in the Spike-driven Transformer. To this end, we design a novel Spike-Driven Self-Attention (SDSA), which exploits only mask and addition operations without any multiplication, and thus having up to $87.2\times$ lower computation energy than vanilla self-attention. Especially in SDSA, the matrix multiplication between Query, Key, and Value is designed as the mask operation. In addition, we rearrange all residual connections in the vanilla Transformer before the activation functions to ensure that all neurons transmit binary spike signals. It is shown that the Spike-driven Transformer can achieve 77.1\% top-1 accuracy on ImageNet-1K, which is the state-of-the-art result in the SNN field. Man Yao, Zhaokun Zhou, Li Yuan 0007, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
NeurIPS | 7 |
| 2023 | Filtered Observations for Model-Based Multi-agent Reinforcement Learning
Linghui Meng 0001, Xuantang Xiong, Yifan Zang 0001, Guoqi Li 0002, Dengpeng Xing, Bo Xu 0002 |
ECML/PKDD (4) | 5 |
| 2023 | Semi-supervised partial label learning algorithm via reliable label propagation
Tian Wang 0001, Guoqi Li 0002, Ming Yan 0007 |
Appl. Intell. | 4 |
| 2023 | Sparser spiking activity can be better: Feature Refine-and-Mask spiking neural network for event-based visual recognition
Man Yao, Hengyu Zhang 0001, Guang-She Zhao, Dingheng Wang, Guoqi Li 0002 |
Neural Networks | 7 |
| 2023 | Attention Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) are becoming a promising energy-efficient alternative to traditional artificial neural networks (ANNs). However, the performance gap between SNNs and ANNs has been a significant hindrance to deploying SNNs ubiquitously. To leverage the full potential of SNNs, in this paper we study the attention mechanisms, which can help human focus on important information. We present our idea of attention in SNNs with a multi-dimensional attention module, which infers attention weights along the temporal, channel, as well as spatial dimension separately or simultaneously. Based on the existing neuroscience theories, we exploit the attention weights to optimize membrane potentials, which in turn regulate the spiking response. Extensive experimental results on event-based action recognition and image classification datasets demonstrate that attention facilitates vanilla SNNs to achieve sparser spiking firing, better performance, and energy efficiency concurrently. In particular, we achieve top-1 accuracy of 75.92% and 77.08% on ImageNet-1 K with single/4-step Res-SNN-104, which are state-of-the-art results in SNNs. Compared with counterpart Res-ANN-104, the performance gap becomes -0.95/+0.21 percent and the energy efficiency is 31.8×/7.4×. To analyze the effectiveness of attention SNNs, we theoretically prove that the spiking degradation or the gradient vanishing, which usually holds in general SNNs, can be resolved by introducing the block dynamical isometry theory. We also analyze the efficiency of attention SNNs based on our proposed spiking response visualization method. Our work lights up SNN's potential as a general backbone to support various applications in the field of SNN research, with a great balance between effectiveness and energy efficiency. Man Yao, Guang-She Zhao, Hengyu Zhang 0001, Yifan Hu 0013, Lei Deng 0003, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Event-Based Semantic Segmentation With Posterior AttentionabstractIn the past years, attention-based Transformers have swept across the field of computer vision, starting a new stage of backbones in semantic segmentation. Nevertheless, semantic segmentation under poor light conditions remains an open problem. Moreover, most papers about semantic segmentation work on images produced by commodity frame-based cameras with a limited framerate, hindering their deployment to auto-driving systems that require instant perception and response at milliseconds. An event camera is a new sensor that generates event data at microseconds and can work in poor light conditions with a high dynamic range. It looks promising to leverage event cameras to enable perception where commodity cameras are incompetent, but algorithms for event data are far from mature. Pioneering researchers stack event data as frames so that event-based segmentation is converted to frame-based segmentation, but characteristics of event data are not explored. Noticing that event data naturally highlight moving objects, we propose a posterior attention module that adjusts the standard attention by the prior knowledge provided by event data. The posterior attention module can be readily plugged into many segmentation backbones. Plugging the posterior attention module into a recently proposed SegFormer network, we get EvSegFormer (the event-based version of SegFormer) with state-of-the-art performance in two datasets (MVSEC and DDD-17) collected for event-based segmentation. Code is available at https://github.com/zexiJia/EvSegFormer to facilitate research on event-based vision. Zexi Jia, Kaichao You, Weihua He, Yang Tian 0002, Yongxiang Feng, Yaoyuan Wang, Xu Jia 0012, Yihang Lou, Guoqi Li 0002 |
IEEE Trans. Image Process. | 10 |
| 2023 | Comprehensive SNN Compression Using ADMM Optimization and Activity RegularizationabstractAs well known, the huge memory and compute costs of both artificial neural networks (ANNs) and spiking neural networks (SNNs) greatly hinder their deployment on edge devices with high efficiency. Model compression has been proposed as a promising technique to improve the running efficiency via parameter and operation reduction, whereas this technique is mainly practiced in ANNs rather than SNNs. It is interesting to answer how much an SNN model can be compressed without compromising its functionality, where two challenges should be addressed: 1) the accuracy of SNNs is usually sensitive to model compression, which requires an accurate compression methodology and 2) the computation of SNNs is event-driven rather than static, which produces an extra compression dimension on dynamic spikes. To this end, we realize a comprehensive SNN compression through three steps. First, we formulate the connection pruning and weight quantization as a constrained optimization problem. Second, we combine spatiotemporal backpropagation (STBP) and alternating direction method of multipliers (ADMMs) to solve the problem with minimum accuracy loss. Third, we further propose activity regularization to reduce the spike events for fewer active operations. These methods can be applied in either a single way for moderate compression or a joint way for aggressive compression. We define several quantitative metrics to evaluate the compression performance for SNNs. Our methodology is validated in pattern recognition tasks over MNIST, N-MNIST, CIFAR10, and CIFAR100 datasets, where extensive comparisons, analyses, and insights are provided. To the best of our knowledge, this is the first work that studies SNN compression in a comprehensive manner by exploiting all compressible components and achieves better results. Lei Deng 0003, Yujie Wu 0002, Yifan Hu 0013, Ling Liang 0003, Guoqi Li 0002, Xing Hu 0001, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Exploring Adversarial Attack in Spiking Neural Networks With Spike-Compatible GradientabstractSpiking neural network (SNN) is broadly deployed in neuromorphic devices to emulate brain function. In this context, SNN security becomes important while lacking in-depth investigation. To this end, we target the adversarial attack against SNNs and identify several challenges distinct from the artificial neural network (ANN) attack: 1) current adversarial attack is mainly based on gradient information that presents in a spatiotemporal pattern in SNNs, hard to obtain with conventional backpropagation algorithms; 2) the continuous gradient of the input is incompatible with the binary spiking input during gradient accumulation, hindering the generation of spike-based adversarial examples; and 3) the input gradient can be all-zeros (i.e., vanishing) sometimes due to the zero-dominant derivative of the firing function. Recently, backpropagation through time (BPTT)-inspired learning algorithms are widely introduced into SNNs to improve the performance, which brings the possibility to attack the models accurately given spatiotemporal gradient maps. We propose two approaches to address the above challenges of gradient-input incompatibility and gradient vanishing. Specifically, we design a gradient-to-spike (G2S) converter to convert continuous gradients to ternary ones compatible with spike inputs. Then, we design a restricted spike flipper (RSF) to construct ternary gradients that can randomly flip the spike inputs with a controllable turnover rate, when meeting all-zero gradients. Putting these methods together, we build an adversarial attack methodology for SNNs. Moreover, we analyze the influence of the training loss function and the firing threshold of the penultimate layer on the attack effectiveness. Extensive experiments are conducted to validate our solution. Besides the quantitative analysis of the influence factors, we also compare SNNs and ANNs against adversarial attacks under different attack methods. This work can help reveal what happens in SNN attacks and might stimulate more research on the security of SNN models and neuromorphic devices. Ling Liang 0003, Xing Hu 0001, Lei Deng 0003, Yujie Wu 0002, Guoqi Li 0002, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Kronecker CP Decomposition With Fast Multiplication for Compressing RNNsabstractRecurrent neural networks (RNNs) are powerful in the tasks oriented to sequential data, such as natural language processing and video recognition. However, because the modern RNNs have complex topologies and expensive space/computation complexity, compressing them becomes a hot and promising topic in recent years. Among plenty of compression methods, tensor decomposition, e.g., tensor train (TT), block term (BT), tensor ring (TR), and hierarchical Tucker (HT), appears to be the most amazing approach because a very high compression ratio might be obtained. Nevertheless, none of these tensor decomposition formats can provide both space and computation efficiency. In this article, we consider to compress RNNs based on a novel Kronecker CANDECOMP/PARAFAC (KCP) decomposition, which is derived from Kronecker tensor (KT) decomposition, by proposing two fast algorithms of multiplication between the input and the tensor-decomposed weight. According to our experiments based on UCF11, Youtube Celebrities Face, UCF50, TIMIT, TED-LIUM, and Spiking Heidelberg digits datasets, it can be verified that the proposed KCP-RNNs have a comparable performance of accuracy with those in other tensor-decomposed formats, and even 278 219× compression ratio could be obtained by the low-rank KCP. More importantly, KCP-RNNs are efficient in both space and computation complexity compared with other tensor-decomposed ones. Besides, we find KCP has the best potential of parallel computing to accelerate the calculations in neural networks. Dingheng Wang, Bijiao Wu, Guang-She Zhao, Man Yao, Hengnu Chen, Lei Deng 0003, Tianyi Yan, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) represent the most prominent biologically inspired computing model for neuromorphic computing (NC) architectures. However, due to the nondifferentiable nature of spiking neuronal functions, the standard error backpropagation algorithm is not directly applicable to SNNs. In this work, we propose a tandem learning framework that consists of an SNN and an artificial neural network (ANN) coupled through weight sharing. The ANN is an auxiliary structure that facilitates the error backpropagation for the training of the SNN at the spike-train level. To this end, we consider the spike count as the discrete neural representation in the SNN and design an ANN neuronal activation function that can effectively approximate the spike count of the coupled SNN. The proposed tandem learning rule demonstrates competitive pattern recognition and regression capabilities on both the conventional frame- and event-based vision datasets, with at least an order of magnitude reduced inference time and total synaptic operations over other state-of-the-art SNN implementations. Therefore, the proposed tandem learning rule offers a novel solution to training efficient, low latency, and high-accuracy deep SNNs with low computing resources. Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li 0002, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelabstractRecently, vision-language pre-training shows great potential in open-vocabulary object detection, where detectors trained on base classes are devised for detecting new classes. The class text embedding is firstly generated by feeding prompts to the text encoder of a pre-trained vision-language model. It is then used as the region classifier to supervise the training of a detector. The key element that leads to the success of this model is the proper prompt, which requires careful words tuning and ingenious design. To avoid laborious prompt engineering, there are some prompt representation learning methods being proposed for the image classification task, which however can only be sub-optimal solutions when applied to the detection task. In this paper, we introduce a novel method, detection prompt (DetPro), to learn continuous prompt representations for open-vocabulary object detection based on the pre-trained vision-language model. Different from the previous classification-oriented methods, DetPro has two highlights: 1) a background interpretation scheme to include the proposals in image background into the prompt training; 2) a context grading scheme to separate proposals in image foreground for tailored prompt training. We assemble DetPro with ViLD, a recent state-of-the-art openworld object detector, and conduct experiments on the LVIS as well as transfer learning on the Pascal VOC, COCO, Objects365 datasets. Experimental results show that our DetPro outperforms the baseline ViLD [7] in all settings, e.g., +3.4 APboxand +3.0 APmaskimprovements on the novel classes of LVIS. Code and models are available at https://github.com/dyabel/detpro. Yu Du 0010, Fangyun Wei, Zihe Zhang, Miaojing Shi, Guoqi Li 0002 |
CVPR | 6 |
| 2022 | Accelerating Spatiotemporal Supervised Training of Large-Scale Spiking Neural Networks on GPUabstractSpiking neural networks (SNNs) have great potential to achieve brain-like intelligence, however, it suffers low accuracy of conventional synaptic plasticity rules and low training efficiency on GPUs. Recently, the emerging backpropagation through time (BPTT) inspired learning algorithms bring new opportunities to boost the accuracy of SNNs, while training on GPUs still remains inefficient due to the complex spatiotemporal dynamics and huge memory consumption, which restricts the model exploration for SNNs and prevents the advance of neuromorphic computing. In this work, we build a framework to solve the inefficiency of BPTT-based SNN training on modern GPUs. To reduce the memory consumption, we optimize the dataflow by saving CONV/FC results only in the forward pass and recomputing other intermediate results in the backward pass. Then, we customize kernel functions to accelerate the neural dynamics for all training stages. Finally, we provide a Pytorch interface to make our framework easy-to-deploy in real systems. Compared to vanilla Pytorch implementation, our framework can achieve up to 2.13 x end-to-end speedup and consume only 0.41 x peak memory on the CIFAR10 dataset. Moreover, for the distributed training on the large ImageNet dataset, we can achieve up to 1.81 x end-to-end speedup and consume only 0.38 x peak memory. Ling Liang 0003, Zhaodong Chen 0001, Lei Deng 0003, Fengbin Tu, Guoqi Li 0002, Yuan Xie 0001 |
DATE | 5 |
| 2022 | Survey on Graph Neural Network Acceleration: An Algorithmic PerspectiveabstractGraph neural networks (GNNs) have been a hot spot of recent research and are widely utilized in diverse applications. However, with the use of huger data and deeper models, an urgent demand is unsurprisingly made to accelerate GNNs for more efficient execution. In this paper, we provide a comprehensive survey on acceleration methods for GNNs from an algorithmic perspective. We first present a new taxonomy to classify existing acceleration methods into five categories. Based on the classification, we systematically discuss these methods and highlight their correlations. Next, we provide comparisons from aspects of the efficiency and characteristics of these methods. Finally, we suggest some promising prospects for future research. Xin Liu 0073, Mingyu Yan, Lei Deng 0003, Guoqi Li 0002, Xiaochun Ye, Dongrui Fan, Shirui Pan, Yuan Xie 0001 |
IJCAI | 4 |
| 2022 | Attention-based Local Mean K-Nearest Centroid Neighbor Classifier
Ming Yan 0007, Guoqi Li 0002, Tian Wang 0001 |
Expert Syst. Appl. | 4 |
| 2022 | Modeling learnable electrical synapse for high precision spatio-temporal recognition
Zhenzhi Wu, Zhihong Zhang 0008, Huanhuan Gao, Rongzhen Zhao, Guang-She Zhao, Guoqi Li 0002 |
Neural Networks | 7 |
| 2022 | A Comprehensive and Modularized Statistical Framework for Gradient Norm Equality in Deep Neural NetworksabstractThe rapid development of deep neural networks (DNNs) in recent years can be attributed to the various techniques that address gradient explosion and vanishing. In order to understand the principle behind these techniques and develop new methods, plenty of metrics have been proposed to identify networks that are free of gradient explosion and vanishing. However, due to the diversity of network components and complex serial-parallel hybrid connections in modern DNNs, the evaluation of existing metrics usually requires strong assumptions, complex statistical analysis, or has limited application fields, which constraints their spread in the community. In this paper, inspired by the Gradient Norm Equality and dynamical isometry, we first propose a novel metric called Block Dynamical Isometry, which measures the change of gradient norm in individual blocks. Because our Block Dynamical Isometry is norm-based, its evaluation needs weaker assumptions compared with the original dynamical isometry. To mitigate challenging derivation, we propose a highly modularized statistical framework based on free probability. Our framework includes several key theorems to handle complex serial-parallel hybrid connections and a library to cover the diversity of network components. Besides, several sufficient conditions for prerequisites are provided. Powered by our metric and framework, we analyze extensive initialization, normalization, and network structures. We find that our Block Dynamical Isometry is a universal philosophy behind them. Then, we improve some existing methods based on our analysis, including an activation function selection strategy for initialization techniques, a new configuration for weight normalization, a depth-aware way to derive coefficients in SeLU, and initialization/weight normalization in DenseNet. Moreover, we propose a novel normalization technique named second moment normalization, which has 30 percent fewer computation overhead than batch normalization without accuracy loss and has better performance under micro batch size. Last but not least, our conclusions and methods are evidenced by extensive experiments on multiple models over CIFAR-10 and ImageNet. Zhaodong Chen 0001, Lei Deng 0003, Bangyan Wang, Guoqi Li 0002, Yuan Xie 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | H2Learn: High-Efficiency Learning Accelerator for High-Accuracy Spiking Neural NetworksabstractAlthough spiking neural networks (SNNs) take benefits from the bioplausible neural modeling, the low accuracy under the common local synaptic plasticity learning rules limits their application in many practical tasks. Recently, an emerging SNN supervised learning algorithm inspired by backpropagation through time (BPTT) from the domain of artificial neural networks (ANNs) has successfully boosted the accuracy of SNNs, and helped improve the practicability of SNNs. However, current general-purpose processors suffer from low efficiency when performing BPTT for SNNs due to the ANN-tailored optimization. On the other hand, current neuromorphic chips cannot support BPTT because they mainly adopt local synaptic plasticity rules for simplified implementation. In this work, we propose H2Learn, a novel architecture that can achieve high efficiency for BPTT-based SNN learning, which ensures high accuracy of SNNs. At the beginning, we characterized the behaviors of BPTT-based SNN learning. Benefited from the binary spike-based computation in the forward pass and weight update, we first design look-up table (LUT)-based processing elements in the forward engine and weight update engine to make accumulations implicit and to fuse the computations of multiple input points. Second, benefited from the rich sparsity in the backward pass, we design a dual-sparsity-aware backward engine, which exploits both input and output sparsity. Finally, we apply a pipeline optimization between different engines to build an end-to-end solution for the BPTT-based SNN learning. Compared with the modern NVIDIA V100 GPU, H2Learn achieves$7.38\times $area saving,$5.74-10.20\times $speedup, and$5.25-7.12\times $energy saving on several benchmark datasets. Ling Liang 0003, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Hardware-Enabled Efficient Data Processing With Tensor-Train DecompositionabstractIn recent years, tensor computation has become a promising tool for solving big data analysis, machine learning, medical image, and EDA problems. To ease the memory and computation intensity of tensor processing, decomposition techniques, especially tensor-train decomposition (TTD), are widely adopted to compress the extremely high-dimensional tensor data. Despite TTD’s potential to break the curse of dimensionality, researchers have not yet leveraged its full computational potential, mainly because of two reasons: 1) executing TTD itself is time- and energy-consuming due to the singular value decomposition (SVD) operation inside each of TTD’s iteration and 2) additional software/hardware optimizations are often required to process the obtained TT-format data in certain applications such as deep learning inference. In this article, we address these challenges with two approaches. First, we propose an algorithm-hardware co-design with customized architecture, namely, TTD Engine to accelerate TTD. We use MRI image compression as a demo application to illustrate the efficacy of the proposed accelerator. Second, we present a case study demonstrating the benefit of TT-format data processing and the efficacy of using TTD Engine. In the case study, we use the TT approach to realize convolution operation, which is difficult and nontrivial for TT-format data. Experimental results show that, TTD Engine achieves, on average,$14.9 \times $–$36.9 \times $speedup over CPU implementations and$4.1\times $–$9.9\times $speedup compared to the GPU baseline. The energy efficiency is also improved by at least$14.4\times $and$5.4\times $over CPU and GPU, respectively. Moreover, our hardware-enabled TT-format data processing further leads to more efficient implementations of complicated operations and applications. Zheng Qu 0002, Lei Deng 0003, Bangyan Wang, Hengnu Chen, Jilan Lin, Ling Liang 0003, Guoqi Li 0002, Zheng Zhang 0005, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | E$^2$ DNet: An Ensembling Deep Neural Network for Solving Nonconvex Economic Dispatch in Smart GridabstractCurrently, a nonconvex economic dispatch problem is one of the research focuses in the field of smart grid (SG). A variety of algorithms are developed to solve it. However, these algorithms are prone to suffering from high computation cost and slow convergence rate, which creates an inevitable gap between theoretical analysis and practical real-time operations. In this article, we aim at providing an ensemble deep-learning-based approach to tackle such a challenging issue. First, a novel ensemble method is presented to explore the ground truth of nonconvex economic dispatch problems. Second, considering the time-varying total load demand, cost coefficients, and dispatchability of all generation units in a practical SG system as the features, a new deep neural network structure is proposed to learn the complex mapping from instant features to an optimal nonconvex economic dispatch solution. If such a mapping is well approximated by the designed deep neural network, no significant effort is required to solve a new economic dispatch problem, and the solution is obtained on the scale of milliseconds. Third, analyzing that a single deep neural network may be weak to a small part of the mapping space of the nonconvex economic dispatch problem, we further present an ensemble of multiple parallel deep neural networks trained sequentially with a simplified Adaboost.R2 algorithm. Finally, case studies reveal that the proposed approach achieves orders of magnitude speedup in computational time while guaranteeing similar or better performance on minimizing the overall generation cost compared to the state-of-the-art nonconvex economic dispatch algorithms. Fanghong Guo, Wen-An Zhang 0001, Guoqi Li 0002, Changyun Wen |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Brain-Controlled 2D Navigation Robot Based on a Spatial Gradient Controller and Predictive Environmental CoordinatorabstractOBJECTIVE: Brain-computer interfaces (BCIs) have been used in two-dimensional (2D) navigation robotic devices, such as brain-controlled wheelchairs and brain-controlled vehicles. However, contemporary BCI systems are driven by binary selective control. On the one hand, only directional information can be transferred from humans to machines, such as "turn left" or "turn right", which means that the quantified value, such as the radius of gyration, cannot be controlled. In this study, we proposed a spatial gradient BCI controller and corresponding environment coordinator, by which the quantified value of brain commands can be transferred in the form of a 2D vector, improving the flexibility, stability and efficiency of BCIs. METHODS: A horizontal array of steady-state visual stimulation was arranged to excite subject (EEG) signals. Covariance arrays between subjects' electroencephalogram (EEG) and stimulation features were mapped into quantified 2-dimensional vectors. The generated vectors were then inputted into the predictive controller and fused with virtual forces generated by the robot's predictive environment coordinator in the form of vector calculation. The resultant vector was then interpreted into the driving force for the robot, and real-time speed feedback was generated. RESULTS: The proposed SGC controller generated a faster (27.4 s vs. 34.9 s) response for the single-obstacle avoidance task than the selective control approach. In practical multiobstacle tasks, the proposed robot executed 39% faster in the target-reaching tasks than the selective controller and had better robustness in multiobstacle avoidance tasks (average failures significantly dropped from 27% to 4%). SIGNIFICANCE: This research proposes a new form of brain-machine shared control strategy that quantifies brain commands in the form of a 2-D control vector stream rather than selective constant values. Combined with a predictive environment coordinator, the brain-controlled strategy of the robot is optimized and provided with higher flexibility. The proposed controller can be used in brain-controlled 2D navigation devices, such as brain-controlled wheelchairs and vehicles. Guoqi Li 0002, Dingjie Suo, Zhiyuan Ming, Tianyi Yan |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | LIAF-Net: Leaky Integrate and Analog Fire Network for Lightweight and Efficient Spatiotemporal Information ProcessingabstractSpiking neural networks (SNNs) based on the leaky integrate and fire (LIF) model have been applied to energy-efficient temporal and spatiotemporal processing tasks. Due to the bioplausible neuronal dynamics and simplicity, LIF-SNN benefits from event-driven processing, however, usually face the embarrassment of reduced performance. This may because, in LIF-SNN, the neurons transmit information via spikes. To address this issue, in this work, we propose a leaky integrate and analog fire (LIAF) neuron model so that analog values can be transmitted among neurons, and a deep network termed LIAF-Net is built on it for efficient spatiotemporal processing. In the temporal domain, LIAF follows the traditional LIF dynamics to maintain its temporal processing capability. In the spatial domain, LIAF is able to integrate spatial information through convolutional integration or fully connected integration. As a spatiotemporal layer, LIAF can also be used with traditional artificial neural network (ANN) layers jointly. In addition, the built network can be trained with backpropagation through time (BPTT) directly, which avoids the performance loss caused by ANN to SNN conversion. Experiment results indicate that LIAF-Net achieves comparable performance to the gated recurrent unit (GRU) and long short-term memory (LSTM) on bAbI question answering (QA) tasks and achieves state-of-the-art performance on spatiotemporal dynamic vision sensor (DVS) data sets, including MNIST-DVS, CIFAR10-DVS, and DVS128 Gesture, with much less number of synaptic weights and computational overhead compared with traditional networks built by LSTM, GRU, convolutional LSTM (ConvLSTM), or 3-D convolution (Conv3D). Compared with traditional LIF-SNN, LIAF-Net also shows dramatic accuracy gain on all these experiments. In conclusion, LIAF-Net provides a framework combining the advantages of both ANNs and SNNs for lightweight and efficient spatiotemporal information processing. Zhenzhi Wu, Hehui Zhang, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Adaptive Control of Second-Order Nonlinear Systems With Injection and Deception AttacksabstractIn this article, the adaptive control for a class of strict-feedback nonlinear systems with uncertainties under injection and deception attacks is considered. An adaptive control scheme is proposed to deal with the injection and deception attacks meanwhile guarantee that regulation errors could be made arbitrarily small by adjusting control parameters. Compared with existing works whose models are linear or relatively simple, the model we consider in this article is nonlinear with parametric uncertainties. A new type of feedback control scheme is introduced to solve this problem. A simulation example is given to verify the effectiveness of our proposed control scheme. Yue Yang 0049, Jiangshuai Huang, Xiaojie Su, Kai Wang 0003, Guoqi Li 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Going Deeper With Directly-Trained Larger Spiking Neural NetworksabstractSpiking neural networks (SNNs) are promising in a bio-plausible coding for spatio-temporal information and event-driven signal processing, which is very suited for energy-efficient implementation in neuromorphic hardware. However, the unique working mode of SNNs makes them more difficult to train than traditional networks. Currently, there are two main routes to explore the training of deep SNNs with high performance. The first is to convert a pre-trained ANN model to its SNN version, which usually requires a long coding window for convergence and cannot exploit the spatio-temporal features during training for solving temporal tasks. The other is to directly train SNNs in the spatio-temporal domain. But due to the binary spike activity of the firing function and the problem of gradient vanishing or explosion, current methods are restricted to shallow architectures and thereby difficult in harnessing large-scale datasets (e.g. ImageNet). To this end, we propose a threshold-dependent batch normalization (tdBN) method based on the emerging spatio-temporal backpropagation, termed “STBP-tdBN”, enabling direct training of a very deep SNN and the efficient implementation of its inference on neuromorphic hardware. With the proposed method and elaborated shortcut connection, we significantly extend directly-trained SNNs from a shallow structure ( Hanle Zheng, Yujie Wu 0002, Lei Deng 0003, Yifan Hu 0013, Guoqi Li 0002 |
AAAI | 5 |
| 2021 | Temporal-wise Attention Spiking Neural Networks for Event Streams ClassificationabstractHow to effectively and efficiently deal with spatio-temporal event streams, where the events are generally sparse and non-uniform and have the μs temporal resolution, is of great value and has various real-life applications. Spiking neural network (SNN), as one of the brain-inspired event-triggered computing models, has the potential to extract effective spatio-temporal features from the event streams. However, when aggregating individual events into frames with a new higher temporal resolution, existing SNN models do not attach importance to that the serial frames have different signal-to-noise ratios since event streams are sparse and non-uniform. This situation interferes with the performance of existing SNNs. In this work, we propose a temporal-wise attention SNN (TA-SNN) model to learn frame-based representation for processing event streams. Concretely, we extend the attention concept to temporal-wise input to judge the significance of frames for the final decision at the training stage, and discard the irrelevant frames at the inference stage. We demonstrate that TA-SNN models improve the accuracy of event streams classification tasks. We also study the impact of multiple-scale temporal resolutions for frame-based representation. Our approach is tested on three different classification tasks: gesture recognition, image classification, and spoken digit recognition. We report the state-of-the-art results on these tasks, and get the essential improvement of accuracy (almost 19%) for gesture recognition with only 60 ms. Man Yao, Huanhuan Gao, Guang-She Zhao, Dingheng Wang, Zhao-Xu Yang, Guoqi Li 0002 |
ICCV | 7 |
| 2021 | Tensor train decomposition for solving large-scale linear equations
Hengnu Chen, Lei Deng 0003, Zheng Qu 0002, Ling Liang 0003, Tianyi Yan, Yuan Xie 0001, Guoqi Li 0002 |
Neurocomputing | 7 |
| 2021 | Training and inference for integer-based semantic segmentation networkabstractSemantic segmentation has been a major topic in research and industry in recent years. However, due to the computation complexity of pixel-wise prediction and backpropagation algorithm, semantic segmentation has been demanding in computation resources, resulting in slow training and inference speed and large storage space to store models. Existing schemes that speed up segmentation network change the network structure and come with noticeable accuracy degradation. However, neural network quantization can be used to reduce computation load while maintaining comparable accuracy and original network structure. Semantic segmentation networks are different from traditional deep convolutional neural networks (DCNNs) in many ways, and this topic has not been thoroughly explored in existing works. In this paper, we propose a new quantization framework for training and inference of segmentation networks, where parameters and operations are constrained to 8-bit integer-based values for the first time. Full quantization of the data flow and the removal of square and root operations in batch normalization give our framework the ability to perform inference on fixed-point devices. Our proposed framework is evaluated on mainstream semantic segmentation networks like FCN-VGG16 and DeepLabv3-ResNet50, achieving comparable accuracy against floating-point framework on ADE20K dataset and PASCAL VOC 2012 dataset. Lei Deng 0003, Yukuan Yang, Yuan Xie 0001, Guoqi Li 0002 |
Neurocomputing | 5 |
| 2021 | QTTNet: Quantized tensor train neural networks for 3D object and video recognition
Donghyun Lee 0002, Dingheng Wang, Yukuan Yang, Lei Deng 0003, Guang-She Zhao, Guoqi Li 0002 |
Neural Networks | 6 |
| 2021 | Nonlinear tensor train format for deep neural network compression
Dingheng Wang, Guang-She Zhao, Hengnu Chen, Zhexian Liu, Lei Deng 0003, Guoqi Li 0002 |
Neural Networks | 6 |
| 2021 | Key Nodes Selection in Controlling Complex Networks via Convex OptimizationabstractKey nodes are the nodes connected with a given number of external source controllers that result in minimal control cost. Finding such a subset of nodes is a challenging task since it impossible to list and evaluate all possible solutions unless the network is small. In this paper, we approximately solve this problem by proposing three algorithms step by step. By relaxing the Boolean constraints in the original optimization model, a convex problem is obtained. Then inexact alternating direction method of multipliers (IADMMs) is proposed and convergence property is theoretically established. Based on the degree distribution, an extension method named degree-based IADMM (D-IADMM) is proposed such that key nodes are pinpointed. In addition, with the technique of local optimization employed on the results of D-IADMM, we also develop LD-IADMM and the performance is greatly improved. The effectiveness of the proposed algorithms is validated on different networks ranging from Erdős-Rényi networks and scale-free networks to some real-life networks. Jie Ding 0007, Changyun Wen, Guoqi Li 0002, Zhenghua Chen |
IEEE Trans. Cybern. | 3 |
| 2021 | Linear Quadratic Optimal Control of Time-Invariant Linear Networks With Selectable Input MatrixabstractOptimal control of networks is to minimize the cost function of a network in a dynamical process with an optimal control strategy. For the time-invariant linear systems, · x(t)=A x(t)+B u(t) , and the traditional linear quadratic regulator (LQR), which minimizes a quadratic cost function, has been well established given both the adjacency matrix A and the control input matrix B . However, this conventional approach is not applicable when we have the freedom to design B . In this article, we investigate the situation when the input matrix B is a variable to be designed to reduce the control cost. First, the problem is formulated and we establish an equivalent expression of the quadratic cost function with respect to B , which is difficult to obtain within the traditional theoretical framework as it requires obtaining an explicit solution of a Riccati differential equation (RDE). Next, we derive the gradient of the quadratic cost function with respect to the matrix variable B analytically. Further, we obtain three inequalities of the cost functions, after which several possible design (optimization) problems are discussed, and algorithms based on gradient information are proposed. It is shown that the cost of controlling the LTI systems can be significantly reduced when the input matrix becomes "designable." We find that the nodes connected to input sources can be sparsely identified and they are distributed as evenly as possible in the LTI networks if one wants to control the networks with the lowest cost. Our findings help us better understand how the LTI systems should be controlled through designing the input matrix. Yukun Hao, Guoqi Li 0002, Changyun Wen |
IEEE Trans. Cybern. | 3 |
| 2021 | Target Controllability of Two-Layer Multiplex Networks Based on Network Flow TheoryabstractIn this paper, we consider the target controllability of two-layer multiplex networks, which is an outstanding challenge faced in various real-world applications. We focus on a fundamental issue regarding how to allocate a minimum number of control sources to guarantee the controllability of each given target subset in each layer, where the external control sources are limited to interact with only one layer. It is shown that this issue is essentially a path cover problem, which is to locate a set of directed paths denoted as P and cycles denoted as C to cover the target sets under the constraint that the nodes in the second layer cannot be the starting node of any element in P , and the number of elements in P attains its minimum. In addition, the formulated path cover problem can be further converted into a maximum network flow problem, which can be efficiently solved by an algorithm called maximum flow-based target path-cover (MFTP). We rigorously prove that MFTP provides the minimum number of control sources for guaranteeing the target controllability of two-layer multiplex networks. It is anticipated that this paper would serve wide applications in target control of real-life networks. Guoqi Li 0002, Xumin Chen, Lei Deng 0003, Gaoxi Xiao, Pei Jing |
IEEE Trans. Cybern. | 2 |
| 2021 | Effective and Efficient Batch Normalization Using a Few Uncorrelated Data for Statistics EstimationabstractDeep neural networks (DNNs) thrive in recent years, wherein batch normalization (BN) plays an indispensable role. However, it has been observed that BN is costly due to the huge reduction and elementwise operations that are hard to be executed in parallel, which heavily reduces the training speed. To address this issue, in this article, we propose a methodology to alleviate the BN's cost by using only a few sampled or generated data for mean and variance estimation at each iteration. The key challenge to reach this goal is how to achieve a satisfactory balance between normalization effectiveness and execution efficiency. We identify that the effectiveness expects less data correlation in sampling while the efficiency expects more regular execution patterns. To this end, we design two categories of approach: sampling or creating a few uncorrelated data for statistics' estimation with certain strategy constraints. The former includes "batch sampling (BS)" that randomly selects a few samples from each batch and "feature sampling (FS)" that randomly selects a small patch from each feature map of all samples, and the latter is "virtual data set normalization (VDN)" that generates a few synthetic random samples to directly create uncorrelated data for statistics' estimation. Accordingly, multiway strategies are designed to reduce the data correlation for accurate estimation and optimize the execution pattern for running acceleration in the meantime. The proposed methods are comprehensively evaluated on various DNN models, where the loss of model accuracy and the convergence rate are negligible. Without the support of any specialized libraries, 1.98× BN layer acceleration and 23.2% overall training speedup can be practically achieved on modern GPUs. Furthermore, our methods demonstrate powerful performance when solving the well-known "micro-BN" problem in the case of a tiny batch size. This article provides a promising solution for the efficient training of high-performance DNNs. Zhaodong Chen 0001, Lei Deng 0003, Guoqi Li 0002, Xing Hu 0001, Ling Liang 0003, Yufei Ding 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Target Controllability in Multilayer Networks via Minimum-Cost Maximum-Flow MethodabstractIn this article, to maximize the dimension of controllable subspace, we consider target controllability problem with maximum covered nodes set in multiplex networks. We call such an issue as maximum-cost target controllability problem. Likewise, minimum-cost target controllability problem is also introduced which is to find minimum covered node set and driver node set. To address these two issues, we first transform them into a minimum-cost maximum-flow problem based on graph theory. Then an algorithm named target minimum-cost maximum-flow (TMM) is proposed. It is shown that the proposed TMM ensures the target nodes in multiplex networks to be controlled with the minimum number of inputs as well as the maximum (minimum) number of covered nodes. Simulation results on Erdős-Rényi (ER-ER) networks, scale-free (SF-SF) networks, and real-life networks illustrate satisfactory performance of the TMM. Jie Ding 0007, Changyun Wen, Guoqi Li 0002, Pengfei Tu, Dongxu Ji, Ying Zou 0002, Jiangshuai Huang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Core Placement Optimization for Multi-chip Many-core Neural Network Systems with Reinforcement LearningabstractMulti-chip many-core neural network systems are capable of providing high parallelism benefited from decentralized execution, and they can be scaled to very large systems with reasonable fabrication costs. As multi-chip many-core systems scale up, communication latency related effects will take a more important portion in the system performance. While previous work mainly focuses on the core placement within a single chip, there are two principal issues still unresolved: the communication-related problems caused by the non-uniform, hierarchical on/off-chip communication capability in multi-chip systems, and the scalability of these heuristic-based approaches in a factorially growing search space. To this end, we propose a reinforcement-learning-based method to automatically optimize core placement through deep deterministic policy gradient, taking into account information of the environment by performing a series of trials (i.e., placements) and using convolutional neural networks to extract spatial features of different placements. Experimental results indicate that compared with a naive sequential placement, the proposed method achieves 1.99× increase in throughput and 50.5% reduction in latency; compared with the simulated annealing, an effective technique to approximate the global optima in an extremely large search space, our method improves the throughput by 1.22× and reduces the latency by 18.6%. We further demonstrate that our proposed method is capable to find optimal placements taking advantages of different communication properties caused by different system configurations, and work in a topology-agnostic manner. Nan Wu 0009, Lei Deng 0003, Guoqi Li 0002, Yuan Xie 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2021 | An Accelerated Distributed Gradient-Based Algorithm for Constrained Optimization With Application to Economic Dispatch in a Large-Scale Power SystemabstractIn this article, we consider a convex optimization problem which minimizes the sum of local agents' cost functions subject to certain local constraints. Besides, both the local cost function and local constraints are only known by the local agent itself. To solve this problem, a new accelerated distributed gradient-based algorithm is proposed, which is inspired by the “momentum” phenomena in nature and aims to accelerate the convergence speed of conventional distributed gradient algorithms. Sufficient conditions for the stepsizes and the acceleration gains are derived to ensure the convergence of the proposed algorithm. Furthermore, based on this proposed fast distributed algorithm, a new decentralized approach is proposed to solve economic dispatch problem, especially for a large-scale power system. Based on the idea of virtual agent, it is proved that this decentralized algorithm is equivalent to the original fast distributed gradient method. Several case studies implemented on IEEE 30-bus, IEEE 118-bus power systems, and a large-scale power system consisting of 1000 generators are conducted to validate the proposed method. Fanghong Guo, Guoqi Li 0002, Changyun Wen, Lei Wang 0059, Ziyang Meng 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | fuseGNN: Accelerating Graph Convolutional Neural Network Training on GPGPUabstractGraph convolutional neural networks (GNN) have achieved state-of-the-art performance on tasks like node classification. It has become a new workload family member in data-centers. GNN works on irregular graph-structured data with three distinct phases: Combination, Graph Processing, and Aggregation. While Combination phase has been well supported by sgemm kernels in cuBLAS, the other two phases are still inefficient on GPGPU due to the lack of optimized CUDA kernels. In particular, Aggregation phase introduces large volume of DRAM storage footprint and data movement, and both Aggregation and Graph Processing phases suffer from high kernel launching time. These inefficiencies not only decrease training throughput but also limit users from training GNNs on larger graphs on GPGPU. Although these problems have been partially alleviated by recent studies, their optimizations are still not sufficient. In this paper, we propose fuseGNN, an extension of PyTorch that provides highly optimized APIs and CUDA kernels for GNN. First, two different programming abstractions for Aggregation phase are utilized to handle graphs with different average degrees. Second, dedicated GPGPU kernels are developed for Aggregation and Graph Processing in both forward and backward passes, in which kernel-fusion along with other optimization strategies are applied to reduce kernel launching time and latency as well as exploit data reuse opportunities. Evaluation on multiple benchmarks shows that fuseGNN achieves up to 5.3× end-to-end speedup over state-of-the-art frameworks, and the DRAM storage footprint is reduced by several orders of magnitude on large datasets. Zhaodong Chen 0001, Mingyu Yan, Maohua Zhu, Lei Deng 0003, Guoqi Li 0002, Shuangchen Li, Yuan Xie 0001 |
ICCAD | 5 |
| 2020 | A deadlock-free physical mapping method on the many-core neural network chip
Guoqi Li 0002, Lei Deng 0003, Guanrui Wang |
Neurocomputing | 3 |
| 2020 | Supervised learning in spiking neural networks with synaptic delay-weight plasticity
Malu Zhang, Jibin Wu, Ammar Belatreche, Zihan Pan, Xiurui Xie, Yansong Chua, Guoqi Li 0002, Hong Qu 0002, Haizhou Li 0001 |
Neurocomputing | 7 |
| 2020 | Parallel alternating direction method of multipliers
Jiaqi Yan 0001, Fanghong Guo, Changyun Wen, Guoqi Li 0002 |
Inf. Sci. | 4 |
| 2020 | Rethinking the performance comparison between SNNS and ANNS
Lei Deng 0003, Yujie Wu 0002, Xing Hu 0001, Ling Liang 0003, Yufei Ding 0001, Guoqi Li 0002, Guang-She Zhao, Peng Li 0001, Yuan Xie 0001 |
Neural Networks | 6 |
| 2020 | Comparing SNNs and RNNs on neuromorphic vision datasets: Similarities and differences
Weihua He, Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Yang Tian 0002, Wenhui Wang 0001, Yuan Xie 0001 |
Neural Networks | 4 |
| 2020 | Compressing 3DCNNs based on tensor train decomposition
Dingheng Wang, Guang-She Zhao, Guoqi Li 0002, Lei Deng 0003, Yang Wu 0001 |
Neural Networks | 3 |
| 2020 | Hybrid tensor decomposition in neural network compression
Bijiao Wu, Dingheng Wang, Guang-She Zhao, Lei Deng 0003, Guoqi Li 0002 |
Neural Networks | 5 |
| 2020 | Training high-performance and large-scale deep neural networks with full 8-bit integers
Yukuan Yang, Lei Deng 0003, Tianyi Yan, Yuan Xie 0001, Guoqi Li 0002 |
Neural Networks | 6 |
| 2020 | Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive SurveyabstractDomain-specific hardware is becoming a promising topic in the backdrop of improvement slow down for general-purpose processors due to the foreseeable end of Moore's Law. Machine learning, especially deep neural networks (DNNs), has become the most dazzling domain witnessing successful applications in a wide spectrum of artificial intelligence (AI) tasks. The incomparable accuracy of DNNs is achieved by paying the cost of hungry memory consumption and high computational complexity, which greatly impedes their deployment in embedded systems. Therefore, the DNN compression concept was naturally proposed and widely used for memory saving and compute acceleration. In the past few years, a tremendous number of compression techniques have sprung up to pursue a satisfactory tradeoff between processing efficiency and application accuracy. Recently, this wave has spread to the design of neural network accelerators for gaining extremely high performance. However, the amount of related works is incredibly huge and the reported approaches are quite divergent. This research chaos motivates us to provide a comprehensive survey on the recent advances toward the goal of efficient compression and execution of DNNs without significantly compromising accuracy, involving both the high-level algorithms and their applications in hardware design. In this article, we review the mainstream compression approaches such as compact model, tensor decomposition, data quantization, and network sparsification. We explain their compression principles, evaluation metrics, sensitivity analysis, and joint-way use. Then, we answer the question of how to leverage these methods in the design of neural network accelerators and present the state-of-the-art hardware architectures. In the end, we discuss several existing issues such as fair comparison, testing workloads, automatic compression, influence on security, and framework/hardware-level support, and give promising topics in this field and the possible challenges as well. This article attempts to enable readers to quickly build up a big picture of neural network compression and acceleration, clearly evaluate various methods, and confidently get started in the right way. Lei Deng 0003, Guoqi Li 0002, Song Han 0003, Luping Shi, Yuan Xie 0001 |
Proc. IEEE | 2 |
| 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on CrossbarsabstractCrossbar architecture has been widely used in neural network (NN) accelerators, involving conventional and emerging devices. It performs well on the fully connected layer through efficient vector-matrix multiplication. Whereas, the advantages degrade on the convolutional layer with huge data reuse, since the execution speed and resource overhead are imbalanced when using existing fully unfolded or fully folded mapping strategy. To address this issue, we propose a novel semi-folded mapping (SemiMap) framework for implementing the convolution on crossbars. It simultaneously folds the physical resources along the row dimension of feature maps (FMs) and unfolds them along the column dimension. The former reduces the resource overhead, and the latter maintains the parallelism. An FM slicing scheme is further proposed to enable the processing of large-size image. Via our mapping framework, a row-by-row streaming pipeline for intraimage dataflow and periodical pipeline for interimage dataflow are easy to be obtained. To validate the idea, we build a many-crossbar architecture with several designs to guarantee the overall functionality and performance. Based on the measurement data of a fabricated chip, a mapping compiler and a cycle-accurate simulator are developed for the hardware simulation of large-scale networks. We evaluate the proposed SemiMap on various convolutional NNs across different network scale. ${>} 35 {\times }$ resource saving and several hundred times cycle reduction are demonstrated compared to the existing fully unfolded and fully folded strategies, respectively. This paper jumps out of the current extreme mapping schemes, and provides a balanced solution on how to efficiently deploy the computational graphs with data reuse on many-crossbar architecture. Lei Deng 0003, Yuan Xie 0001, Ling Liang 0003, Guanrui Wang, Liang Chang 0002, Xing Hu 0001, Liu Liu 0017, Jing Pei, Guoqi Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2020 | Automatic Cataract Classification Using Deep Neural Network With Discrete State TransitionabstractCataract is the clouding of lens, which affects vision and it is the leading cause of blindness in the world's population. Accurate and convenient cataract detection and cataract severity evaluation will improve the situation. Automatic cataract detection and grading methods are proposed in this paper. With prior knowledge, the improved Haar features and visible structure features are combined as features, and multilayer perceptron with discrete state transition (DST-MLP) or exponential DST (EDST-MLP) are designed as classifiers. Without prior knowledge, residual neural networks with DST (DST-ResNet) or EDST (EDST-ResNet) are proposed. Whether with prior knowledge or not, our proposed DST and EDST strategy can prevent overfitting and reduce storage memory during network training and implementation, and neural networks with these strategies achieve state-of-the-art accuracy in cataract detection and grading. The experimental results indicate that combined features always achieve better performance than a single type of feature, and classification methods with feature extraction based on prior knowledge are more suitable for complicated medical image classification task. These analyses can provide constructive advice for other medical image processing applications. Yue Zhou 0014, Guoqi Li 0002, Huiqi Li |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Matrix Function Optimization Problems Under Orthonormal ConstraintabstractWe investigate the matrix function optimization under the orthonormal constraint on the matrix variable. By introducing an index-notation-arrangement-based chain rule (I-Chain rule), we obtain the gradient of the cost function and propose a revisited orthonormal-constraint-based projected gradient method to locate a minimum of an objective/cost function of matrix variables iteratively subject to orthonormal constraint. To guarantee the convergence the proposed method, existing schemes require the gradient can be represented by the multiplication of a symmetrical matrix and the matrix variable itself. This condition has been relaxed in this paper. New techniques are proposed to establish the convergence property of the iterative algorithm. Simulation results show the effectiveness of our framework. This paper allows more extensive applications of matrix function optimization problems in science and engineering. Guoqi Li 0002, Huiqi Li, A. K. Qin 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Direct Training for Spiking Neural Networks: Faster, Larger, BetterabstractSpiking neural networks (SNNs) that enables energy efficient implementation on emerging neuromorphic hardware are gaining more attention. Yet now, SNNs have not shown competitive performance compared with artificial neural networks (ANNs), due to the lack of effective learning algorithms and efficient programming frameworks. We address this issue from two aspects: (1) We propose a neuron normalization technique to adjust the neural selectivity and develop a direct learning algorithm for deep SNNs. (2) Via narrowing the rate coding window and converting the leaky integrate-and-fire (LIF) model into an explicitly iterative version, we present a Pytorch-based implementation method towards the training of large-scale SNNs. In this way, we are able to train deep SNNs with tens of times speedup. As a result, we achieve significantly better accuracy than the reported works on neuromorphic datasets (N-MNIST and DVSCIFAR10), and comparable accuracy as existing ANNs and pre-trained SNNs on non-spiking datasets (CIFAR10). To our best knowledge, this is the first work that demonstrates direct training of deep SNNs with high performance on CIFAR10, and the efficient implementation provides a new way to explore the potential of SNNs. Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Jun Zhu 0001, Yuan Xie 0001, Luping Shi |
AAAI | 3 |
| 2019 | Dynamic Sparse Graph for Efficient Deep Learning
Liu Liu 0017, Lei Deng 0003, Xing Hu 0001, Maohua Zhu, Guoqi Li 0002, Yufei Ding 0001, Yuan Xie 0001 |
ICLR (Poster) | 5 |
| 2019 | Deep Spiking Neural Network with Spike Count based Learning RuleabstractDeep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training. Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001 |
IJCNN | 5 |
| 2019 | Minimum Cost Control of Directed Networks With Selectable Control InputsabstractThe minimum cost control problem is one of the most important issues in controlling complex networks. Different from the previous works, in this paper, we consider the minimum cost control problem with selectable inputs by adopting the cost function summed over both quadratic terms of system input and system state with a weighting factor. To address such an issue, the orthonormal-constraint-based projected gradient method is proposed to determine the input matrix iteratively. Convergence of the proposed algorithm is established. Extensive simulation results are carried out to show the effectiveness of the proposed algorithm. We also investigate what kinds of nodes are most important for minimizing average control cost in directed stems/circles and small networks through simulation studies. The presented results in this paper bring meaningful physical insights in controlling the directed networks from an energy point of view. Guoqi Li 0002, Jie Ding 0007, Changyun Wen, Jiangshuai Huang |
IEEE Trans. Cybern. | 1 |
| 2019 | L1-Norm Batch Normalization for Efficient Training of Deep Neural NetworksabstractBatch normalization (BN) has recently become a standard component for accelerating and improving the training of deep neural networks (DNNs). However, BN brings in additional calculations, consumes more memory, and significantly slows down the training iteration. Furthermore, the nonlinear square and sqrt operations in the normalization process impede low bit-width quantization techniques, which draw much attention to the deep learning hardware community. In this paper, we propose an$L1$-norm BN (L1BN) with only linear operations in both forward and backward propagations during training. L1BN is approximately equivalent to the conventional$L2$-norm BN (L2BN) by multiplying a scaling factor that equals$({\pi }/{2})^{1/2}$. Experiments on various convolutional neural networks and generative adversarial networks reveal that L1BN can maintain the same performance and convergence rate as L2BN but with higher computational efficiency. In real application-specified integrated circuit synthesis with reduced resources, L1BN achieves 25% speedup and 37% energy saving compared to the original L2BN. Our hardware-friendly normalization method not only surpasses L2BN in speed but also simplifies the design of deep learning accelerators. Last but not least, L1BN promises a fully quantized training of DNNs, which empowers future artificial intelligence applications on mobile devices with transfer and continual learning capability. Guoqi Li 0002, Lei Deng 0003, Liu Liu 0017, Yuan Xie 0001, Luping Shi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | HitNet: Hybrid Ternary Recurrent Neural NetworkabstractQuantization is a promising technique to reduce the model size, memory footprint, and massive computation operations of recurrent neural networks (RNNs) for embedded devices with limited resources. Although extreme low-bit quantization has achieved impressive success on convolutional neural networks, it still suffers from huge accuracy degradation on RNNs with the same low-bit precision. In this paper, we first investigate the accuracy degradation on RNN models under different quantization schemes, and the distribution of tensor values in the full precision model. Our observation reveals that due to the difference between the distributions of weights and activations, different quantization methods are suitable for different parts of models. Based on our observation, we propose HitNet, a hybrid ternary recurrent neural network, which bridges the accuracy gap between the full precision model and the quantized model. In HitNet, we develop a hybrid quantization method to quantize weights and activations. Moreover, we introduce a sloping factor motivated by prior work on Boltzmann machine to activation functions, further closing the accuracy gap between the full precision model and the quantized model. Overall, our HitNet can quantize RNN models into ternary values, {-1, 0, 1}, outperforming the state-of-the-art quantization methods on RNN models significantly. We test it on typical RNN models, such as Long-Short-Term Memory (LSTM) and Gated Recurrent Units (GRU), on which the results outperform previous work significantly. For example, we improve the perplexity per word (PPW) of a ternary LSTM on Penn Tree Bank (PTB) corpus from 126 (the state-of-the-art result to the best of our knowledge) to 110.3 with a full precision model in 97.2, and a ternary GRU from 142 to 113.5 with a full precision model in 102.7. Peiqi Wang 0001, Xinfeng Xie, Lei Deng 0003, Guoqi Li 0002, Dongsheng Wang 0002, Yuan Xie 0001 |
NeurIPS | 4 |
| 2018 | Fully Distributed Adaptive Consensus Control of a Class of High-Order Nonlinear Systems With a Directed Topology and Unknown Control DirectionsabstractIn this paper, we investigate the adaptive consensus control for a class of high-order nonlinear systems with different unknown control directions where communications among the agents are represented by a directed graph. Based on backstepping technique, a fully distributed adaptive control approach is proposed without using global information of the topology. Meanwhile, a novel Nussbaum-type function is proposed to address the consensus control with unknown control directions. It is proved that boundedness of all closed-loop signals and asymptotically consensus tracking for all the agents' outputs are ensured. In simulation studies, a numerical example is illustrated to show the effectiveness of the control scheme. Jiangshuai Huang, Yongduan Song 0001, Wei Wang 0016, Changyun Wen, Guoqi Li 0002 |
IEEE Trans. Cybern. | 5 |
| 2017 | Energy Generation Scheduling in Microgrids Involving Temporal-Correlated Renewable EnergyabstractIn this paper, a cost minimization problem is formulated to intelligently schedule energy generations for microgrids equipped with unstable renewable sources and energy storages. In such systems, the uncertain renewable energy will impose unprecedented scheduling challenges. To cope with the fluctuate nature of the renewable energy, an uncertainty model based on renewable energies' moment statistics is developed. Specifically, we obtain the mean vector and second-order moment matrix according to predictions and field measurements and then define uncertainty set to confine the renewable energy generation. The uncertainty model allows the renewable energy generation distributions to fluctuate within the uncertainty set. We develop chance constraint approximations and robust optimization approaches based on a Chebyshev inequality framework to firstly transform and then solve the scheduling problem. Numerical results based on real-world data traces evaluate the performance bounds of the proposed scheduling scheme. It is shown that the temporal-correlation information of the renewable energy within a proper time span can effectively reduce the conservativeness of the solution. Moreover, detailed studies on the impacts of different factors on the proposed scheme provide some interesting insights which shall be useful for the policy making for the future microgrids. Ran Wang 0004, Gaoxi Xiao, Ping Wang 0001, Yue Cao 0002, Guoqi Li 0002, Jie Hao 0002, Kun Zhu 0001 |
GLOBECOM | 5 |
| 2017 | Boundary Constraints for Minimum Cost Control of Directed NetworksabstractControlling directed networks with minimum cost has become an emerging branch in the areas of complex networks and control recently. In this paper, we focus on this minimum cost control problem subject to two types of boundary constraints, namely, trace boundary constraint and orthonormal boundary constraint on the input matrices. First, the minimum cost control problem is formulated as an optimization model for each type of boundary constraint. Next, two iterative algorithms, named as trace-constraint-based projected gradient method and orthonormal-constraint-based projected gradient method, are proposed to solve the optimal problem, respectively. Then, convergence properties of both algorithms are established. Finally, extensive simulation results show the effectiveness of our methods based on detailed comparisons between the two boundary conditions. We believe the results reveal some interesting physical insights for the optimal control of directed networks. Guoqi Li 0002, Pei Tang, Changyun Wen, Ziyang Meng 0001 |
IEEE Trans. Cybern. | 1 |
| 2016 | Locality sensitive batch feature extraction for high-dimensional data
Jie Ding 0007, Changyun Wen, Guoqi Li 0002, Chin-Seng Chua |
Neurocomputing | 3 |
| 2014 | Min-max discriminant analysis based on gradient method for feature extractionabstractFeature extraction is an essential step in pattern classification, which is normally divided into two tasks: transforming the input vector into a feature vector and/or reducing its dimensionality. A well-defined feature extraction algorithm makes the subsequent classification process more effective and efficient. One of the most important feature extraction algorithms is linear discriminant analysis (LDA). However, there is a critical drawback for LDA. For a classification task with c classes, since the rank of the between class matrix cannot be larger than c - 1, the dimension of the projected subspace is at most c - 1 for LDA. From this viewpoint, min-max discriminant analysis based on gradient method (MMDA-GM) is derived in this paper. With the proposed MMDA-GM, a set of features can be extracted simultaneously. It is shown that the proposed method achieves good performance for data sets from UCI Machine Learning Repository. Jie Ding 0007, Guoqi Li 0002, Changyun Wen, Chin-Seng Chua |
ICARCV | 2 |
| 2014 | Support vector machine based liver cancer early detection using magnetic resonance imagesabstractMagnetic Resonance Imaging (MRI) has become an important tool for doctors to diagnose liver cancer for decays. The survival rate of liver cancer patients can be significantly improved by an early diagnosis. In this paper, we present a computer aided kernel based support vector machine (SVM) algorithm for diagnosing liver cancer in early stage by applying our proposed method to the patients' magnetic resonance (MR) images. We apply the histogram-based feature extraction method to extract feature information from each raw MR image acquired. And 100 confirmed liver cancer and 100 confirmed benign type liver tumor (BLT) patients' feature information are used to form our training data set to train or SVM classification engine. The model is tested with a set of 30 confirmed early stage liver cancer and 30 BLT samples. Our trained SVM achieves an accuracy of 86.67% in classifying early stage liver cancer and 80.00% in classifying BLT. Lei Meng 0001, Changyun Wen, Guoqi Li 0002 |
ICARCV | 3 |
| 2013 | Revised online learning with kernels for classification and regressionabstractRevised algorithm for online learning with kernels (OLK) in classification and regression is proposed in a reproducing kernel hilbert space (RKHS). Compared with the original OLK, the revised algorithm allows that the new data points arrive either one by one or two by two. Guoqi Li 0002, Ning Ning 0001, Kiruthika Ramanathan, Luping Shi |
CIDM | 1 |
| 2013 | Behind the magical numbers: Hierarchical Chunking and the Human Working Memory CapacityabstractTo explore the influence of chunking on the capacity limits of working memory, a model for chunking in sequential working memory is proposed, using hierarchical bidirectional inhibition-connected neural networks with winnerless competition. With the assumption of the existence of an upper bound to the inhibitory weights in neurobiological networks, it is shown that chunking increases the number of memorized items in working memory from the "magical number 7" to 16 items. The optimal number of chunks and the number of the memorized items in each chunk are the "magical number 4". Guoqi Li 0002, Ning Ning 0001, Kiruthika Ramanathan, Luping Shi |
Int. J. Neural Syst. | 1 |
| 2013 | Model-Based Online Learning With KernelsabstractNew optimization models and algorithms for online learning with Kernels (OLK) in classification, regression, and novelty detection are proposed in a reproducing Kernel Hilbert space. Unlike the stochastic gradient descent algorithm, called the naive online Reg minimization algorithm (NORMA), OLK algorithms are obtained by solving a constrained optimization problem based on the proposed models. By exploiting the techniques of the Lagrange dual problem like Vapnik's support vector machine (SVM), the solution of the optimization problem can be obtained iteratively and the iteration process is similar to that of the NORMA. This further strengthens the foundation of OLK and enriches the research area of SVM. We also apply the obtained OLK algorithms to problems in classification, regression, and novelty detection, including real time background substraction, to show their effectiveness. It is illustrated that, based on the experimental results of both classification and regression, the accuracy of OLK algorithms is comparable with traditional SVM-based algorithms, such as SVM and least square SVM (LS-SVM), and with the state-of-the-art algorithms, such as Kernel recursive least square (KRLS) method and projectron method, while it is slightly higher than that of NORMA. On the other hand, the computational cost of the OLK algorithm is comparable with or slightly lower than existing online methods, such as above mentioned NORMA, KRLS, and projectron methods, but much lower than that of SVM-based algorithms. In addition, different from SVM and LS-SVM, it is possible for OLK algorithms to be applied to non-stationary problems. Also, the applicability of OLK in novelty detection is illustrated by simulation results. Guoqi Li 0002, Changyun Wen, Zhengguo Li, Feng Yang 0011, Kezhi Mao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Presynaptic Learning and Memory with a Persistent Firing Neuron and a Habituating Synapse: a Model of Short Term Persistent HabituationabstractOur paper explores the interaction of persistent firing axonal and presynaptic processes in the generation of short term memory for habituation. We first propose a model of a sensory neuron whose axon is able to switch between passive conduction and persistent firing states, thereby triggering short term retention to the stimulus. Then we propose a model of a habituating synapse and explore all nine of the behavioral characteristics of short term habituation in a two neuron circuit. We couple the persistent firing neuron to the habituation synapse and investigate the behavior of short term retention of habituating response. Simulations show that, depending on the amount of synaptic resources, persistent firing either results in continued habituation or maintains the response, both leading to longer recovery times. The effectiveness of the model as an element in a bio-inspired memory system is discussed. Kiruthika Ramanathan, Ning Ning 0001, Dhiviya Dhanasekar, Guoqi Li 0002, Luping Shi, Prahlad Vadakkepat |
Int. J. Neural Syst. | 4 |
| 2011 | Error tolerance based support vector machine for regression
Guoqi Li 0002, Changyun Wen, Guang-Bin Huang |
Neurocomputing | 1 |
| 2010 | Identification of Wiener systems based on fixed point theoryabstractIn this paper, we propose a new method for the identification of Wiener systems based on fixed point theory. The linear part of the system is an infinite impulse response (IIR) system and the nonlinear static function is allowed to be non-continuous or non-smooth. Our proposed technique transforms the estimation of parameters to finding a fixed point of a nonlinear equation. We show the existence of the fixed point and also develop an iterative algorithm to find the fixed point. It is proved that, the determined fixed point is actually a global minimum point of the cost function and it is unique, and thus global convergence of the estimates is ensured. The performance of the proposed approach is illustrated by simulation studies. Guoqi Li 0002, Changyun Wen |
ICARCV | 1 |