VLDB 2026 Research / reviewers in the wild / expert
Xuerui Qiu
dblp:351/8271
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Deep learning architectures and training · 36% Efficient and distributed learning · 29% Trustworthy machine learning · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Emerging computing paradigms · 95% Energy-efficient computing · 5% | |
| Network and information security
2 papers |
Digital forensics and information hiding · 100% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
spiking neural network |
3.5 | 5 | 2025 | Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025 High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024 |
Emerging computing paradigms
neuromorphic computing |
2.8 | 4 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025 Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025 |
Digital forensics and information hiding
watermarking |
1.7 | 2 | 2025 | Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025 WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025 |
Emerging computing paradigms › neuromorphic computing
spiking neural network |
1.7 | 2 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning · AAAI 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient neural network training |
0.9 | 1 | 2025 | Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 |
Computer vision › 3D vision › 3d object recognition
point cloud recognition |
0.9 | 1 | 2025 | Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.9 | 1 | 2025 | Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025 |
Digital forensics and information hiding › watermarking › multimedia watermarking
video watermarking |
0.9 | 1 | 2025 | Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking · NeurIPS 2025 |
Emerging computing paradigms › neuromorphic computing › spiking neural network
spiking transformer |
0.9 | 1 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.8 | 1 | 2024 | RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding · ACM Multimedia 2024 |
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient neural network |
0.8 | 1 | 2024 | Gated Attention Coding for Training High-Performance and Efficient Spiking Neural Networks · AAAI 2024 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
0.8 | 1 | 2024 | High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.8 | 1 | 2024 | High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
randomized smoothing |
0.8 | 1 | 2024 | RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding · ACM Multimedia 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2025 | WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models · ICML 2025 |
Energy-efficient computing › energy-efficient machine learning
energy-efficient neural network inference |
0.3 | 1 | 2025 | Quantized Spike-driven Transformer · ICLR 2025 |
Computer vision › Video understanding and tracking
temporal modeling |
0.2 | 1 | 2024 | High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
three-stage learning · 1.7spike voxel coding · 1.7spike sparse convolution · 1.7pseudo-ensemble training · 1.7mutual information · 1.7multimodal large language model · 1.7knowledge distillation · 1.7event-driven learning · 1.7attention zoneout · 1.7visual question answering · 0.9state space model · 0.9mamba · 0.9bilevel optimization · 0.9bi-level optimization · 0.93d wavelet transform · 0.9observer model analysis · 0.8gated attention · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient 3D Recognition with Event-driven Spike Sparse ConvolutionabstractSpiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature. Xuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou, Shibo Zhou, Bo Xu 0002, Guoqi Li 0002 |
AAAI | 1 |
| 2025 | Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction LearningabstractRecent advancements in neuroscience research have propelled the development of Spiking Neural Networks (SNNs), which not only have the potential to further advance neuroscience research but also serve as an energy-efficient alternative to Artificial Neural Networks (ANNs) due to their spike-driven characteristics. However, previous studies often overlooked the multiscale information and its spatiotemporal correlation between event data, leading SNN models to approximate each frame of input events as static images. We hypothesize that this oversimplification significantly contributes to the performance gap between SNNs and traditional ANNs. To address this issue, we have designed a Spiking Multiscale Attention (SMA) module that captures multiscale spatiotemporal interaction information. Furthermore, we developed a regularization method named Attention ZoneOut (AZO), which utilizes spatiotemporal attention weights to reduce the model's generalization error through pseudo-ensemble training. Our approach has achieved state-of-the-art results on mainstream neuromorphic datasets. Additionally, we have reached a performance of 77.1\% on the Imagenet-1K dataset using a 104-layer ResNet architecture enhanced with SMA and AZO. This achievement confirms the state-of-the-art performance of SNNs with non-transformer architectures and underscores the effectiveness of our method in bridging the performance gap between SNN models and traditional ANN models. Yimeng Shan, Malu Zhang, Rui-Jie Zhu 0003, Xuerui Qiu, Jason Kamran Eshraghian, Haicheng Qu |
AAAI | 4 |
| 2025 | Quantized Spike-driven TransformerabstractSpiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm.
However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer structures, which typically rely on substantial computational resources, limiting their deployment on resource-constrained devices.
To overcome this challenge, we propose a quantized spike-driven Transformer baseline (QSD-Transformer), which achieves reduced resource demands by utilizing a low bit-width parameter.
Regrettably, the QSD-Transformer often suffers from severe performance degradation.
In this paper, we first conduct empirical analysis and find that the bimodal distribution of quantized spike-driven self-attention (Q-SDSA) leads to spike information distortion (SID) during quantization, causing significant performance degradation. To mitigate this issue, we take inspiration from mutual information entropy and propose a bi-level optimization strategy to rectify the information distribution in Q-SDSA.
Specifically, at the lower level, we introduce an information-enhanced LIF to rectify the information distribution in Q-SDSA.
At the upper level, we propose a fine-grained distillation scheme for the QSD-Transformer to align the distribution in Q-SDSA with that in the counterpart ANN.
By integrating the bi-level optimization strategy, the QSD-Transformer can attain enhanced energy efficiency without sacrificing its high-performance advantage.
We validate the QSD-Transformer on various visual tasks, and experimental results indicate that our method achieves state-of-the-art results in the SNN domain.
For instance, when compared to the prior SNN benchmark on ImageNet, the QSD-Transformer achieves 80.3\% top-1 accuracy, accompanied by significant reductions of 6.0$\times$ and 8.1$\times$ in power consumption and model size, respectively. Code is available at https://github.com/bollossom/QSD-Transformer. Xuerui Qiu, Malu Zhang, Jieyuan Zhang, Wenjie Wei, Honglin Cao, Junsheng Guo, Rui-Jie Zhu 0003, Yimeng Shan, Yang Yang 0002, Haizhou Li 0001 |
ICLR | 1 |
| 2025 | WMarkGPT: Watermarked Image Understanding via Multimodal Large Language ModelsabstractInvisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in text-driven generative watermarking and fail to capture critical aspects of watermarking, particularly visibility. More importantly, these metrics fail to account for potential corruption of image content. To address these limitations, we propose WMarkGPT, the first multimodal large language model (MLLM) specifically designed for comprehensive watermarked image understanding, without accessing original images. WMarkGPT not only predicts watermark visibility but also generates detailed textual descriptions of its location, content, and impact on image semantics, enabling a more nuanced interpretation of watermarked images. Tackling the challenge of precise location description and understanding images with vastly different content, we construct three visual question-answering (VQA) datasets: an object location-aware dataset, a synthetic watermarking dataset, and a real watermarking dataset. We introduce a meticulously designed three-stage learning pipeline to progressively equip WMarkGPT with the necessary abilities. Extensive experiments on synthetic and real watermarking QA datasets demonstrate that WMarkGPT outperforms existing MLLMs, achieving significant improvements in visibility prediction and content description. The datasets and code are released at https://github.com/TanSongBai/WMarkGPT. Songbai Tan, Xuerui Qiu, Yao Shu, Linrui Xu, Huiping Zhuang, Ming Li 0011, F. Richard Yu |
ICML | 2 |
| 2025 | Safe-Sora: Safe Text-to-Video Generation via Graphical WatermarkingabstractThe explosive growth of generative video models has amplified the demand for
reliable copyright preservation of AI-generated content. Despite its popularity in
image synthesis, invisible generative watermarking remains largely underexplored
in video generation. To address this gap, we propose Safe-Sora, the first framework
to embed graphical watermarks directly into the video generation process. Motivated by the observation that watermarking performance is closely tied to the visual
similarity between the watermark and cover content, we introduce a hierarchical
coarse-to-fine adaptive matching mechanism. Specifically, the watermark image is
divided into patches, each assigned to the most visually similar video frame, and
further localized to the optimal spatial region for seamless embedding. To enable
spatiotemporal fusion of watermark patches across video frames, we develop a 3D
wavelet transform-enhanced Mamba architecture with a novel scanning strategy,
effectively modeling long-range dependencies during watermark embedding and
retrieval. To the best of our knowledge, this is the first attempt to apply state space
models to watermarking, opening new avenues for efficient and robust watermark
protection. Extensive experiments demonstrate that Safe-Sora achieves state-of-the-
art performance in terms of video quality, watermark fidelity, and robustness, which
is largely attributed to our proposals. Code and additional supporting materials are
provided in the supplementary. Zihan Su, Xuerui Qiu, Tangyu Jiang, Junhao Zhuang, Chun Yuan 0003, Ming Li 0073, Shengfeng He, F. Richard Yu |
NeurIPS | 2 |
| 2025 | Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation TrainingabstractThe ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5%, 79.8%, 84.0%, and 86.2% with models containing 10 M, 19 M, 83 M, and 173 M parameters, respectively. For instance, the 10 M model outperforms the best existing SNN by 7.2% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5× and 3.9×, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone. Man Yao, Xuerui Qiu, Tianxiang Hu, Yuhong Chou, Keyu Tian, Jianxing Liao, Luziwei Leng, Bo Xu 0002, Guoqi Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Gated Attention Coding for Training High-Performance and Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are emerging as an energy-efficient alternative to traditional artificial neural networks (ANNs) due to their unique spike-based event-driven nature. Coding is crucial in SNNs as it converts external input stimuli into spatio-temporal feature sequences. However, most existing deep SNNs rely on direct coding that generates powerless spike representation and lacks the temporal dynamics inherent in human vision. Hence, we introduce Gated Attention Coding (GAC), a plug-and-play module that leverages the multi-dimensional gated attention unit to efficiently encode inputs into powerful representations before feeding them into the SNN architecture. GAC functions as a preprocessing layer that does not disrupt the spike-driven nature of the SNN, making it amenable to efficient neuromorphic hardware implementation with minimal modifications. Through an observer model theoretical analysis, we demonstrate GAC's attention mechanism improves temporal dynamics and coding efficiency. Experiments on CIFAR10/100 and ImageNet datasets demonstrate that GAC achieves state-of-the-art accuracy with remarkable efficiency. Notably, we improve top-1 accuracy by 3.10% on CIFAR100 with only 6-time steps and 1.07% on ImageNet while reducing energy usage to 66.9% of the previous works. To our best knowledge, it is the first time to explore the attention-based dynamic coding scheme in deep SNNs, with exceptional effectiveness and efficiency on large-scale datasets. Code is available at https://github.com/bollossom/GAC. Xuerui Qiu, Rui-Jie Zhu 0003, Yuhong Chou, Zhaorui Wang 0005, Liang-Jian Deng, Guoqi Li 0002 |
AAAI | 1 |
| 2024 | High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost
Man Yao, Xuerui Qiu, Yuhong Chou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
ICML | 3 |
| 2024 | RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
Keming Wu, Man Yao, Yuhong Chou, Xuerui Qiu, Bo Xu 0002, Guoqi Li 0002 |
ACM Multimedia | 4 |
| 2024 | Tensor decomposition based attention module for spiking neural networks
Rui-Jie Zhu 0003, Xuerui Qiu, Yule Duan 0001, Malu Zhang, Liang-Jian Deng |
Knowl. Based Syst. | 3 |