Xuerui Qiu

dblp:351/8271 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Deep learning architectures and training · 36% Efficient and distributed learning · 29% Trustworthy machine learning · 12%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Emerging computing paradigms · 95% Energy-efficient computing · 5%
Network and information security
2 papers
Digital forensics and information hiding · 100%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
spiking neural network
3.552025
Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024
Emerging computing paradigms
neuromorphic computing
2.842025
Quantized Spike-driven Transformer · ICLR 2025
Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
Digital forensics and information hiding
watermarking
1.722025
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025
WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025
Emerging computing paradigms › neuromorphic computing
spiking neural network
1.722025
Quantized Spike-driven Transformer · ICLR 2025
Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025
Machine learning › Generative modeling
diffusion model
0.912025
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient neural network training
0.912025
Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Quantized Spike-driven Transformer · ICLR 2025
Computer vision › 3D vision › 3d object recognition
point cloud recognition
0.912025
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Quantized Spike-driven Transformer · ICLR 2025
Machine learning › Generative modeling › video generation
text-to-video generation
0.912025
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025
Digital forensics and information hiding › watermarking › multimedia watermarking
video watermarking
0.912025
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025
Emerging computing paradigms › neuromorphic computing › spiking neural network
spiking transformer
0.912025
Quantized Spike-driven Transformer · ICLR 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.812024
RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding · ACM Multimedia 2024
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient neural network
0.812024
Gated Attention Coding for Training High-Performance and Efficient Spiking Neural Networks · AAAI 2024
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network
0.812024
High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024
Machine learning › Efficient and distributed learning
memory-efficient training
0.812024
High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024
Machine learning › Trustworthy machine learning › robustness › certified robustness
randomized smoothing
0.812024
RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding · ACM Multimedia 2024
Machine learning › Trustworthy machine learning
interpretability
0.312025
WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025
Computer vision › Vision and language
visual question answering
0.312025
WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025
Energy-efficient computing › energy-efficient machine learning
energy-efficient neural network inference
0.312025
Quantized Spike-driven Transformer · ICLR 2025
Computer vision › Video understanding and tracking
temporal modeling
0.212024
High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024

Methods — techniques the papers use, named apart from their topics

three-stage learning · 1.7spike voxel coding · 1.7spike sparse convolution · 1.7pseudo-ensemble training · 1.7mutual information · 1.7multimodal large language model · 1.7knowledge distillation · 1.7event-driven learning · 1.7attention zoneout · 1.7visual question answering · 0.9state space model · 0.9mamba · 0.9bilevel optimization · 0.9bi-level optimization · 0.93d wavelet transform · 0.9observer model analysis · 0.8gated attention · 0.8
YearPublicationVenuePosition
2025 Efficient 3D Recognition with Event-driven Spike Sparse Convolution
abstract
Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature.
Xuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou, Shibo Zhou, Bo Xu 0002, Guoqi Li 0002
AAAI1
2025 Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning
abstract
Recent advancements in neuroscience research have propelled the development of Spiking Neural Networks (SNNs), which not only have the potential to further advance neuroscience research but also serve as an energy-efficient alternative to Artificial Neural Networks (ANNs) due to their spike-driven characteristics. However, previous studies often overlooked the multiscale information and its spatiotemporal correlation between event data, leading SNN models to approximate each frame of input events as static images. We hypothesize that this oversimplification significantly contributes to the performance gap between SNNs and traditional ANNs. To address this issue, we have designed a Spiking Multiscale Attention (SMA) module that captures multiscale spatiotemporal interaction information. Furthermore, we developed a regularization method named Attention ZoneOut (AZO), which utilizes spatiotemporal attention weights to reduce the model's generalization error through pseudo-ensemble training. Our approach has achieved state-of-the-art results on mainstream neuromorphic datasets. Additionally, we have reached a performance of 77.1\% on the Imagenet-1K dataset using a 104-layer ResNet architecture enhanced with SMA and AZO. This achievement confirms the state-of-the-art performance of SNNs with non-transformer architectures and underscores the effectiveness of our method in bridging the performance gap between SNN models and traditional ANN models.
Yimeng Shan, Malu Zhang, Rui-Jie Zhu 0003, Xuerui Qiu, Jason Kamran Eshraghian, Haicheng Qu
AAAI4
2025 Quantized Spike-driven Transformer
abstract
Spiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm. However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer structures, which typically rely on substantial computational resources, limiting their deployment on resource-constrained devices. To overcome this challenge, we propose a quantized spike-driven Transformer baseline (QSD-Transformer), which achieves reduced resource demands by utilizing a low bit-width parameter. Regrettably, the QSD-Transformer often suffers from severe performance degradation. In this paper, we first conduct empirical analysis and find that the bimodal distribution of quantized spike-driven self-attention (Q-SDSA) leads to spike information distortion (SID) during quantization, causing significant performance degradation. To mitigate this issue, we take inspiration from mutual information entropy and propose a bi-level optimization strategy to rectify the information distribution in Q-SDSA. Specifically, at the lower level, we introduce an information-enhanced LIF to rectify the information distribution in Q-SDSA. At the upper level, we propose a fine-grained distillation scheme for the QSD-Transformer to align the distribution in Q-SDSA with that in the counterpart ANN. By integrating the bi-level optimization strategy, the QSD-Transformer can attain enhanced energy efficiency without sacrificing its high-performance advantage. We validate the QSD-Transformer on various visual tasks, and experimental results indicate that our method achieves state-of-the-art results in the SNN domain. For instance, when compared to the prior SNN benchmark on ImageNet, the QSD-Transformer achieves 80.3\% top-1 accuracy, accompanied by significant reductions of 6.0$\times$ and 8.1$\times$ in power consumption and model size, respectively. Code is available at https://github.com/bollossom/QSD-Transformer.
Xuerui Qiu, Malu Zhang, Jieyuan Zhang, Wenjie Wei, Honglin Cao, Junsheng Guo, Rui-Jie Zhu 0003, Yimeng Shan, Yang Yang 0002, Haizhou Li 0001
ICLR1
2025 WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models
abstract
Invisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in text-driven generative watermarking and fail to capture critical aspects of watermarking, particularly visibility. More importantly, these metrics fail to account for potential corruption of image content. To address these limitations, we propose WMarkGPT, the first multimodal large language model (MLLM) specifically designed for comprehensive watermarked image understanding, without accessing original images. WMarkGPT not only predicts watermark visibility but also generates detailed textual descriptions of its location, content, and impact on image semantics, enabling a more nuanced interpretation of watermarked images. Tackling the challenge of precise location description and understanding images with vastly different content, we construct three visual question-answering (VQA) datasets: an object location-aware dataset, a synthetic watermarking dataset, and a real watermarking dataset. We introduce a meticulously designed three-stage learning pipeline to progressively equip WMarkGPT with the necessary abilities. Extensive experiments on synthetic and real watermarking QA datasets demonstrate that WMarkGPT outperforms existing MLLMs, achieving significant improvements in visibility prediction and content description. The datasets and code are released at https://github.com/TanSongBai/WMarkGPT.
Songbai Tan, Xuerui Qiu, Yao Shu, Linrui Xu, Huiping Zhuang, Ming Li 0011, F. Richard Yu
ICML2
2025 Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
abstract
The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-Sora, the first framework to embed graphical watermarks directly into the video generation process. Motivated by the observation that watermarking performance is closely tied to the visual similarity between the watermark and cover content, we introduce a hierarchical coarse-to-fine adaptive matching mechanism. Specifically, the watermark image is divided into patches, each assigned to the most visually similar video frame, and further localized to the optimal spatial region for seamless embedding. To enable spatiotemporal fusion of watermark patches across video frames, we develop a 3D wavelet transform-enhanced Mamba architecture with a novel scanning strategy, effectively modeling long-range dependencies during watermark embedding and retrieval. To the best of our knowledge, this is the first attempt to apply state space models to watermarking, opening new avenues for efficient and robust watermark protection. Extensive experiments demonstrate that Safe-Sora achieves state-of-the- art performance in terms of video quality, watermark fidelity, and robustness, which is largely attributed to our proposals. Code and additional supporting materials are provided in the supplementary.
Zihan Su, Xuerui Qiu, Tangyu Jiang, Junhao Zhuang, Chun Yuan 0003, Ming Li 0073, Shengfeng He, F. Richard Yu
NeurIPS2
2025 Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training
abstract
The ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5%, 79.8%, 84.0%, and 86.2% with models containing 10 M, 19 M, 83 M, and 173 M parameters, respectively. For instance, the 10 M model outperforms the best existing SNN by 7.2% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5× and 3.9×, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone.
Man Yao, Xuerui Qiu, Tianxiang Hu, Yuhong Chou, Keyu Tian, Jianxing Liao, Luziwei Leng, Bo Xu 0002, Guoqi Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Gated Attention Coding for Training High-Performance and Efficient Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are emerging as an energy-efficient alternative to traditional artificial neural networks (ANNs) due to their unique spike-based event-driven nature. Coding is crucial in SNNs as it converts external input stimuli into spatio-temporal feature sequences. However, most existing deep SNNs rely on direct coding that generates powerless spike representation and lacks the temporal dynamics inherent in human vision. Hence, we introduce Gated Attention Coding (GAC), a plug-and-play module that leverages the multi-dimensional gated attention unit to efficiently encode inputs into powerful representations before feeding them into the SNN architecture. GAC functions as a preprocessing layer that does not disrupt the spike-driven nature of the SNN, making it amenable to efficient neuromorphic hardware implementation with minimal modifications. Through an observer model theoretical analysis, we demonstrate GAC's attention mechanism improves temporal dynamics and coding efficiency. Experiments on CIFAR10/100 and ImageNet datasets demonstrate that GAC achieves state-of-the-art accuracy with remarkable efficiency. Notably, we improve top-1 accuracy by 3.10% on CIFAR100 with only 6-time steps and 1.07% on ImageNet while reducing energy usage to 66.9% of the previous works. To our best knowledge, it is the first time to explore the attention-based dynamic coding scheme in deep SNNs, with exceptional effectiveness and efficiency on large-scale datasets. Code is available at https://github.com/bollossom/GAC.
Xuerui Qiu, Rui-Jie Zhu 0003, Yuhong Chou, Zhaorui Wang 0005, Liang-Jian Deng, Guoqi Li 0002
AAAI1
2024 High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost
Man Yao, Xuerui Qiu, Yuhong Chou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002
ICML3
2024 RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
Keming Wu, Man Yao, Yuhong Chou, Xuerui Qiu, Bo Xu 0002, Guoqi Li 0002
ACM Multimedia4
2024 Tensor decomposition based attention module for spiking neural networks
Rui-Jie Zhu 0003, Xuerui Qiu, Yule Duan 0001, Malu Zhang, Liang-Jian Deng
Knowl. Based Syst.3