Rong Xiao 0001

dblp:75/5560-1 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-6408-9724ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Transfer learning and domain adaptation · 34% Efficient and distributed learning · 25% Deep learning architectures and training · 24%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
1.422026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Fast and Accurate Classification with a Multi-Spike Learning Algorithm for Spiking Neurons · IJCAI 2019
Machine learning › Deep learning architectures and training
spiking neural network
1.022023
Towards Energy-Preserving Natural Language Understanding With Spiking Neural Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2023
STCA: Spatio-Temporal Credit Assignment with Delayed Feedback in Deep Spiking Neural Networks · IJCAI 2019
Machine learning › Efficient and distributed learning › model compression
sparse neural network
1.012026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Emerging computing paradigms › neuromorphic computing › neuromorphic vision
event-based vision
1.012026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Machine learning › Transfer learning and domain adaptation › zero-shot learning
generalized zero-shot learning
0.912025
Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective · IEEE Trans. Image Process. 2025
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding
0.912025
Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective · IEEE Trans. Image Process. 2025
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.912025
Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective · IEEE Trans. Image Process. 2025
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.412019
Fast and Accurate Classification with a Multi-Spike Learning Algorithm for Spiking Neurons · IJCAI 2019
Machine learning › Efficient and distributed learning
model compression
0.312026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Information retrieval › image retrieval › semantic image retrieval
zero-shot image retrieval
0.312025
Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective · IEEE Trans. Image Process. 2025

Methods — techniques the papers use, named apart from their topics

sparse encoding · 2.0dual-branch architecture · 2.0asynchronous processing · 2.0vision transformer · 1.7topological alignment · 1.7per-instance attribute synthesis · 1.7cosine similarity calibration · 1.7spiking encoder · 0.7bi-directional spiking neural network · 0.7spatiotemporal error backpropagation · 0.4multi-spike learning rule · 0.4membrane potential trace · 0.4
YearPublicationVenuePosition
2026 S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing
abstract
Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels.
Rong Xiao 0001, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang
AAAI2
2026 MSER: Multi-scale event representation model for enhanced spatio-temporal feature extraction
Wanying Xu, Rong Xiao 0001, Chenwei Tang, Jiancheng Lv 0001, Huajin Tang
Neurocomputing4
2026 IF4FD: Multiscale Information Fusion for Zero-Shot Industrial Fault Diagnosis
abstract
Fault diagnosis aims to identify faults occurring in industrial production processes to prevent personnel injuries and economic losses. However, there are two main challenges in solving fault diagnosis, i.e.,extracting discriminative features from limited sensor dataandrecognizing new classes of faults. To fill these research gaps, we propose a zero-shot fault diagnosis framework, calledIF4FD, based on multiscale information fusion. First, we enhanced raw data from the perspectives of category knowledge, attribute knowledge, and feature knowledge. Then, by drawing on zero-shot learning (ZSL), we can transfer knowledge of trained faults to new classes of faults, enabling the classification of previously unknown faults. The multiscale informative knowledge effectively facilitates knowledge transfer and fault classification, thereby enhancing the accuracy of zero-shot fault diagnosis. Extensive experiments on two industrial fault diagnosis datasets validate the effectiveness of the proposed method, which consistently achieves superior performance compared to representative zero-shot fault diagnosis methods, general ZSL baselines, and several supervised classifiers. A case study on real industrial data from the Cranfield Multiphase Flow Facility also confirms the method’s effectiveness in practical applications.
Chenwei Tang, Wangyang Ying, Nanxu Gong, Wei Ju 0001, Rong Xiao 0001, Jiancheng Lv 0001
IEEE Trans. Ind. Informatics9
2025 SaENeRF: Suppressing Artifacts in Event-based Neural Radiance Fields
abstract
Event cameras are neuromorphic vision sensors that asynchronously capture changes in logarithmic brightness changes, offering significant advantages such as low latency, low power consumption, low bandwidth, and high dynamic range. While these characteristics make them ideal for high-speed scenarios, reconstructing geometrically consistent and photometrically accurate 3D representations from event data remains fundamentally challenging. Current event-based Neural Radiance Fields (NeRF) methods partially address these challenges but suffer from persistent artifacts caused by aggressive network learning in early stages and the inherent noise of event cameras. To overcome these limitations, we present SaENeRF, a novel self-supervised framework that effectively suppresses artifacts and enables 3D-consistent, dense, and photorealistic NeRF reconstruction of static scenes solely from event streams. Our approach normalizes predicted radiance variations based on accumulated event polarities, facilitating progressive and rapid learning for scene representation construction. Additionally, we introduce regularization losses specifically designed to suppress artifacts in regions where photometric changes fall below the event threshold and simultaneously enhance the light intensity difference of non-zero events, thereby improving the visual fidelity of the reconstructed scene. Extensive qualitative and quantitative experiments demonstrate that our method significantly reduces artifacts and achieves superior reconstruction quality compared to existing methods. The code is available at https://github.com/Mr-firework/SaENeRF.
Yuanjian Wang, Yufei Deng, Rong Xiao 0001, Chenwei Tang, Deng Xiong, Jiancheng Lv 0001
IJCNN3
2025 Multi-attribute dynamic attenuation learning improved spiking actor network
Rong Xiao 0001, Jie Zhang 0012, Chenwei Tang, Jiancheng Lv 0001
Neurocomputing1
2025 Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective
abstract
Generalized zero-shot learning (GZSL) shows great potential for improving generalization to unseen classes in real-world scenarios. However, most GZSL methods depend on benchmark datasets with per-class attribute annotations, which creates a large semantic gap and worsens the domain shift problem in the visual-semantic space. To address these challenges, instance-level attributes offer an intuitive solution, but they require expensive manual annotation. In this paper, we propose a simple yet effective approach called per-instance attribute synthesis (PIAS) to generate diverse semantic representations for each instance. Our method first uses the Vision Transformer (ViT) model to extract visual features and then generates per-instance attributes. The patch splitting, positional embedding, and multi-head self-attention mechanisms in ViT improve the discriminability of both visual and semantic representations. Next, we define the generated attributes of class-average images as class anchor points. These anchor points are calibrated in the semantic space by minimizing the cosine similarity between the anchor points and per-class attribute annotations. Finally, we improve the diversity of generated per-instance attributes by aligning the topological structure between per-class attribute annotations and synthesized per-instance attributes with that between class-average visual features and per-instance visual features. We conduct comprehensive experiments on three challenging ZSL datasets: AWA2, CUB, and SUN. The results show that PIAS significantly outperforms state-of-the-art methods under both ZSL and GZSL settings. We further demonstrate the generalization ability of PIAS by applying it to attribute-based zero-shot image retrieval tasks.
Chenwei Tang, Qianjun Zhang, Rong Xiao 0001, Zhenan He 0001, Jiancheng Lv 0001
IEEE Trans. Image Process.5
2025 STSF: Spiking Time Sparse Feedback Learning for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are biologically plausible models known for their computational efficiency. A significant advantage of SNNs lies in the binary information transmission through spike trains, eliminating the need for multiplication operations. However, due to the spatio-temporal nature of SNNs, direct application of traditional backpropagation (BP) training still results in significant computational costs. Meanwhile, learning methods based on unsupervised synaptic plasticity provide an alternative for training SNNs but often yield suboptimal results. Thus, efficiently training high-accuracy SNNs remains a challenge. In this article, we propose a highly efficient and biologically plausible spiking time sparse feedback (STSF) learning method. This algorithm modifies synaptic weights by incorporating a neuromodulator for global supervised learning using sparse direct feedback alignment (DFA) and local homeostasis learning with vanilla spike-timing-dependent plasticity (STDP). Such neuromorphic global-local learning focuses on instantaneous synaptic activity, enabling independent and simultaneous optimization of each network layer, thereby improving biological plausibility, enhancing parallelism, and reducing storage overhead. Incorporating sparse fixed random feedback connections for global error modulation, which uses selection operations instead of multiplication operations, further improves computational efficiency. Experimental results demonstrate that the proposed algorithm markedly reduces the computational cost with significantly higher accuracy comparable to current state-of-the-art algorithms across a wide range of classification tasks. Our implementation codes are available at https://github.com/hppeace/STSF.
Rong Xiao 0001, Chenwei Tang, Shudong Huang, Jiancheng Lv 0001, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.2
2024 Improving generalized zero-shot learning via cluster-based semantic disentangling representation
Wentao Feng, Rong Xiao 0001, Lihuo He, Zhenan He 0001, Jiancheng Lv 0001, Chenwei Tang
Pattern Recognit.3
2023 Towards Energy-Preserving Natural Language Understanding With Spiking Neural Networks
abstract
Artificial neural networks have shown promising results in a variety of natural language understanding (NLU) tasks. Despite their successes, conventional neural-based NLU models are criticized for high energy consumption, making them laborious to be widely applied in low-power electronics, such as smartphones and intelligent terminals. In this paper, we introduce a potential direction to alleviate this bottleneck by proposing a spiking encoder. The core of our model is bi-directional spiking neural network (SNN) which transforms numeric values into discrete spiking signals and replaces massive multiplications with much cheaper additive operations. We examine our model on sentiment classification and machine translation tasks. Experimental results reveal that our model achieves comparable classification and translation accuracy to advancedTransformerbaseline, whereas significantly reduces the required computational energy to 0.82%.
Rong Xiao 0001, Yu Wan 0004, Baosong Yang, Haibo Zhang 0013, Huajin Tang, Derek F. Wong, Boxing Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Event stream learning using spatio-temporal event surface
Junfei Dong, Runhao Jiang, Rong Xiao 0001, Rui Yan 0005, Huajin Tang
Neural Networks3
2020 A Supervised Learning Algorithm for Learning Precise Timing of Multispike in Multilayer Spiking Neural Networks
Rong Xiao 0001, Tianyu Geng
ICONIP (5)1
2020 An Event-Driven Categorization Model for AER Image Sensors Using Multispike Encoding and Learning
abstract
In this article, we present a systematic computational model to explore brain-based computation for object recognition. The model extracts temporal features embedded in address-event representation (AER) data and discriminates different objects by using spiking neural networks (SNNs). We use multispike encoding to extract temporal features contained in the AER data. These temporal patterns are then learned through the tempotron learning rule. The presented model is consistently implemented in a temporal learning framework, where the precise timing of spikes is considered in the feature-encoding and learning process. A noise-reduction method is also proposed by calculating the correlation of an event with the surrounding spatial neighborhood based on the recently proposed time-surface technique. The model evaluated on wide spectrum data sets (MNIST, N-MNIST, MNIST-DVS, AER Posture, and Poker Card) demonstrates its superior recognition performance, especially for the events with noise.
Rong Xiao 0001, Huajin Tang, Rui Yan 0005, Garrick Orchard
IEEE Trans. Neural Networks Learn. Syst.1
2019 STCA: Spatio-Temporal Credit Assignment with Delayed Feedback in Deep Spiking Neural Networks
abstract
The temporal credit assignment problem, which aims to discover the predictive features hidden in distracting background streams with delayed feedback, remains a core challenge in biological and machine learning. To address this issue, we propose a novel spatio-temporal credit assignment algorithm called STCA for training deep spiking neural networks (DSNNs). We present a new spatiotemporal error backpropagation policy by defining a temporal based loss function, which is able to credit the network losses to spatial and temporal domains simultaneously. Experimental results on MNIST dataset and a music dataset (MedleyDB) demonstrate that STCA can achieve comparable performance with other state-of-the-art algorithms with simpler architectures. Furthermore, STCA successfully discovers predictive sensory features and shows the highest performance in the unsegmented sensory event detection tasks.
Pengjie Gu, Rong Xiao 0001, Gang Pan 0001, Huajin Tang
IJCAI2
2019 Fast and Accurate Classification with a Multi-Spike Learning Algorithm for Spiking Neurons
abstract
The formulation of efficient supervised learning algorithms for spiking neurons is complicated and remains challenging. Most existing learning methods with the precisely firing times of spikes often result in relatively low efficiency and poor robustness to noise. To address these limitations, we propose a simple and effective multi-spike learning rule to train neurons to match their output spike number with a desired one. The proposed method will quickly find a local maximum value (directly related to the embedded feature) as the relevant signal for synaptic updates based on membrane potential trace of a neuron, and constructs an error function defined as the difference between the local maximum membrane potential and the firing threshold. With the presented rule, a single neuron can be trained to learn multi-category tasks, and can successfully mitigate the impact of the input noise and discover embedded features. Experimental results show the proposed algorithm has higher precision, lower computation cost, and better noise robustness than current state-of-the-art learning methods under a wide range of learning tasks.
Rong Xiao 0001, Qiang Yu 0005, Rui Yan 0005, Huajin Tang
IJCAI1
2018 Spike-based encoding and learning of spectrum features for robust sound recognition
Rong Xiao 0001, Huajin Tang, Pengjie Gu
Neurocomputing1
2017 An Event-Driven Computational System with Spiking Neurons for Object Recognition
Rong Xiao 0001, Huajin Tang
ICONIP (6)2