EDBT 2026 Demo / reviewers in the wild / expert
Yichen Xiao
dblp:260/2179
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantization-aware distributed deep reinforcement learning for dynamic multi-robot scheduling
Yichen Xiao, Kaixin Cui, Dawei Shi |
Expert Syst. Appl. | 2 |
| 2026 | FDAU-Net: An efficient optic disc segmentation network with flexible wavelet convolution block and selective dual-attention fusion gateabstractAccurate segmentation of optic disc regions in fundus images is essential for the diagnosis and monitoring ocular fundus pathology of ophthalmic diseases such as myopia, glaucoma, and diabetic retinopathy. However, challenges such as irregular disc shapes, edge detail preservation, and limited annotated datasets hinder the effectiveness of existing methods. To handle these problems, we propose FDAU-Net, an efficient, fully convolutional network for optic disc segmentation. FDAU-Net integrates the Flexible Wavelet Convolution Block (FWC Block), which combines the dynamic receptive field adjustment of AKConv with wavelet down-sampling to achieve efficient feature extraction while retaining critical edge information. Additionally, we design the Selective Dual-Attention Fusion Gate (SDA Gate) to optimize feature selection and fusion, leveraging channel and spatial attention mechanisms to substantially enhance segmentation accuracy and robustness. To further improve performance, we introduce MedAugment, an automated data augmentation method tailored for medical images, efficiently boosts data diversity in scenarios with low datasets. Experimental results on the iChallenge and IDRiD datasets demonstrate that FDAU-Net consistently achieves superior segmentation performance across key metrics, highlighting its potential for advancing clinical applications. Yichen Xiao, Ziwei Xiang, Teruko Fukuyama, Yanze Yu, Shengtao Liu, Wuzhida Bao, Shiping Wen 0001, Xingtao Zhou |
Neural Networks | 1 |
| 2025 | Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual ProcessingabstractEvent cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utilize conventional frame-based computer vision algorithms. These frame-based methods, though typically effective, diminish the inherent advantages of the event camera's low latency and low power consumption. To solve the above problems, we propose ASGCN, which efficiently processes data on an event-by-event basis and dynamically evolves into a corresponding dynamic representation, enabling low latency and high sparsity of data representation. The sparsity computation is further improved by introducing brain-inspired spiking neural networks, resulting in low power consumption for ASGCN. Extensive and diverse experiments demonstrate the energy efficiency and low latency advantages of our processing pipeline. Especially on real-world event camera datasets, our pipeline consumes more than 10,000 times less energy and achieves similar performance compared to current frame-based methods. Dingyi Zeng, Honglin Cao, Wanlong Liu, Yichen Xiao, Chengzhuo Lu, Wenyu Chen 0001, Malu Zhang, Guoqing Wang 0001, Yang Yang 0002 |
AAAI | 5 |
| 2025 | Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking TransformersabstractTransformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyze the cause of this gap and attribute it to the dot product’s ineffectiveness in measuring similarity between spiking queries and keys, due to numerous non-spiking events. To address this, we propose a novel α-XNOR similarity measure tailored for spike trains. It redefines the correlation between non-spike pairs as a specific value α, effectively overcoming the limitations of dot-product similarity. Furthermore, considering the sparse nature of spike trains where spikes carry more information than non-spikes, the α-XNOR similarity correspondingly highlights the distinct importance of spikes over non-spikes. Extensive experiments demonstrate that α-XNOR similarity significantly improves performance across different spiking Transformer architectures on various static and neuromorphic datasets, further revealing the potential of spiking Transformers. Yichen Xiao, Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Yimeng Shan, Yulin Jiang, Malu Zhang |
CVPR | 1 |
| 2025 | Mixed-Precision Graph Neural Quantization for Low Bit Large Language ModelsabstractPost-Training Quantization (PTQ) is pivotal for deploying large language models (LLMs) within resource-limited settings by significantly reducing resource demands. However, existing PTQ strategies underperform at low bit levels (< 3 bits) due to the significant difference between the quantized and original weights. To enhance the quantization performance at low bit widths, we introduce a Mixed-precision Graph Neural PTQ (MG-PTQ) approach, employing a graph neural network (GNN) module to capture dependencies among weights and adaptively assign quantization bit-widths. Through the information propagation of the GNN module, our method more effectively captures dependencies among target weights, leading to a more accurate assessment of weight importance and optimized allocation of quantization strategies. Extensive experiments on the WikiText2 and C4 datasets demonstrate that our MG-PTQ method outperforms previous state-of-the-art PTQ method GPTQ, setting new benchmarks for quantization performance under low-bit (< 3 bits) conditions. Wanlong Liu, Yichen Xiao, Dingyi Zeng, Hongyang Zhao, Wenyu Chen 0001, Malu Zhang |
ICASSP | 2 |
| 2025 | Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction ModelingabstractDocument-level relation extraction aims at extracting relational facts between two entities in a document. Existing approaches mainly focus on target entities, utilizing techniques such as graph neural networks to enhance their representations. However, they ignore the rich semantic correlations among entity pairs which provide wider and multifaceted information at a higher level. In this paper, we propose the Relation-based Entity-pair-level Inference (REI) model, which facilitates information interaction at the entity-pair level, enhancing logical reasoning among entities and capturing semantic correlations among entity pairs. Our REI model comprises two modules: Relation-based Information Aggregation (RIA) and Entity-pair-level Information Interaction (EII). The RIA module builds and integrates relation representations to filter out distractions from unrelated entity pairs, while the EII module models entity-pair-level information interaction through multi-head attentions. Extensive experiments on the DocRED, DWIE, CDR, and GDA datasets demonstrate the superiority of the proposed REI model, outperforming previous state-of-the-art approaches. Furthermore, we provide detailed experimental analyses based on the performance gains and illustrate the interpretability. Wanlong Liu, Dingyi Zeng, Li Zhou 0010, Yichen Xiao, Malu Zhang, Wenyu Chen 0001 |
ICASSP | 4 |
| 2025 | Memory-Free and Parallel Computation for Quantized Spiking Neural NetworksabstractQuantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause for this decline: the loss of historical information due to the quantized membrane potential. To tackle this issue, we introduce a memory-free quantization method that captures all historical information without directly storing membrane potentials, resulting in better performance with less memory requirements. To further improve the computational efficiency, we propose a parallel training and asynchronous inference framework that greatly increases training speed and energy efficiency. We combine the proposed memory-free quantization and parallel computation methods to develop a high-performance and efficient QSNN, named MFP-QSNN. Extensive experiments show that our MFP-QSNN achieves state-of-the-art performance on various static and neuromorphic image datasets, requiring less memory and faster training speeds. The efficiency and efficacy of the MFP-QSNN highlight its potential for energy-efficient neuromorphic computing. Dehao Zhang, Shuai Wang 0058, Yichen Xiao, Wenjie Wei, Yimeng Shan, Malu Zhang, Yang Yang 0002 |
ICASSP | 3 |
| 2025 | Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group TransformationsabstractAdapting pre-trained foundation models for diverse downstream tasks is a core practice in artificial intelligence. However, the wide range of tasks and high computational costs make full fine-tuning impractical. To overcome this, parameter-efficient fine-tuning (PEFT) methods like LoRA have emerged and are becoming a growing research focus. Despite the success of these methods, they are primarily designed for linear layers, focusing on two-dimensional matrices while largely ignoring higher-dimensional parameter spaces like convolutional kernels. Moreover, directly applying these methods to higher-dimensional parameter spaces often disrupts their structural relationships. Given the rapid advancements in matrix-based PEFT methods, rather than designing a specialized strategy, we propose a generalization that extends matrix-based PEFT methods to higher-dimensional parameter spaces without compromising their structural properties. Specifically, we treat parameters as elements of a Lie group, with updates modeled as perturbations in the corresponding Lie algebra. These perturbations are mapped back to the Lie group through the exponential map, ensuring smooth, consistent updates that preserve the inherent structure of the parameter space. Extensive experiments on computer vision and natural language processing validate the effectiveness and versatility of our approach, demonstrating clear improvements over existing methods. Chongjie Si, Zhiyi Shi, Yichen Xiao, Xiaokang Yang 0001, Wei Shen 0002 |
ICCV | 4 |
| 2025 | Spiking Vision Transformer with Saccadic AttentionabstractThe combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN counterparts. Here, we first analyze why SNN-based ViTs suffer from limited performance and identify a mismatch between the vanilla self-attention mechanism and spatio-temporal spike trains. This mismatch results in degraded spatial relevance and limited temporal interactions. To address these issues, we draw inspiration from biological saccadic attention mechanisms and introduce an innovative Saccadic Spike Self-Attention (SSSA) method. Specifically, in the spatial domain, SSSA employs a novel spike distribution-based method to effectively assess the relevance between Query and Key pairs in SNN-based ViTs. Temporally, SSSA employs a saccadic interaction module that dynamically focuses on selected visual areas at each timestep and significantly enhances whole scene understanding through temporal interactions.
Building on the SSSA mechanism, we develop a SNN-based Vision Transformer (SNN-ViT). Extensive experiments across various visual tasks demonstrate that SNN-ViT achieves state-of-the-art performance with linear computational complexity. The effectiveness and efficiency of the SNN-ViT highlight its potential for power-critical edge vision applications. Shuai Wang 0058, Malu Zhang, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Yimeng Shan, Qian Sun 0014, Enqi Zhang, Yang Yang 0002 |
ICLR | 5 |
| 2025 | Bipolar Self-attention for Spiking TransformersabstractHarnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores.
To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation.
Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers. Shuai Wang 0058, Malu Zhang, Dehao Zhang, Yimeng Shan, Jieyuan Zhang, Yichen Xiao, Honglin Cao, Zeyu Ma 0002, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 7 |
| 2025 | Document-level relation extraction with structural encoding and entity-pair-level information interaction
Wanlong Liu, Yichen Xiao, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Weishan Kong, Malu Zhang, Wenyu Chen 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Efficient Automatic Modulation Classification in Nonterrestrial Networks With SNN-Based TransformerabstractWith the development of informatization of IoT devices, nonterrestrial networks (NTNs) are becoming more and more important. NTN, including air and space networks, face challenges, such as high-computational complexity, bandwidth requirements, and memory constraints. An intelligent automatic modulation classification (AMC) mechanism based on neural networks plays a pivotal role in enhancing spectrum efficiency, throughput, and link reliability. Past work in AMC has evolved from likelihood-based and feature-based methods to traditional machine learning techniques and, more recently, to deep neural networks (DNNs). However, existing DNN architectures pose challenges for NTN due to high-computational complexity, bandwidth requirements, and memory consumption. Addressing this problems, we proposes a spiking transformer-based model for AMC, exploiting temporal dynamics for enhanced performance. Biologically inspired spiking neural networks enable us to exploit the sparse and binarized activation properties of spiking neurons, allowing us to build AMC models with high-energy efficiency and high availability that can be used in NTN systems. Furthermore, we introduce a weight binarization method to reduce the model size, which also further reduces the bandwidth and memory requirements of AMC in NTN edge deployment. Experimental results demonstrate the superiority of our approach over state-of-the-art methods, with the binarized model achieving comparable accuracy at a fraction of the size. Dingyi Zeng, Yichen Xiao, Wanlong Liu, Huilin Du, Enqi Zhang, Dehao Zhang, Malu Zhang, Wenyu Chen 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Ternary spike-based neuromorphic signal processing system
Shuai Wang 0058, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Hongyu Qing, Wenjie Wei, Malu Zhang, Yang Yang 0002 |
Neural Networks | 4 |
| 2025 | MIU-Net: Advanced multi-scale feature extraction and imbalance mitigation for optic disc segmentationabstractPathological myopia is a severe eye condition that can cause serious complications like retinal detachment and macular degeneration, posing a threat to vision. Optic disc segmentation helps measure changes in the optic disc and observe the surrounding retina, aiding early detection of pathological myopia. However, these changes make segmentation difficult, resulting in accuracy levels that are not suitable for clinical use. To address this, we propose a new model called MIU-Net, which improves segmentation performance through several innovations. First, we introduce a multi-scale feature extraction (MFE) module to capture features at different scales, helping the model better identify optic disc boundaries in complex images. Second, we design a dual attention module that combines channel and spatial attention to focus on important features and improve feature use. To tackle the imbalance between optic disc and background pixels, we use focal loss to enhance the model's ability to detect minority optic disc pixels. We also apply data augmentation techniques to increase data diversity and address the lack of training data. Our model was tested on the iChallenge-PM and iChallenge-AMD datasets, showing clear improvements in accuracy and robustness compared to existing methods. The experimental results demonstrate the effectiveness and potential of our model in diagnosing pathological myopia and other medical image processing tasks. Yichen Xiao, Shengtao Liu, Teruko Fukuyama, Xiaoliao Peng, Guangyang Tian, Shiping Wen 0001, Xingtao Zhou |
Neural Networks | 1 |
| 2024 | Q-SNNs: Quantized Spiking Neural NetworksabstractBrain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to represent information and process them in an asynchronous event-driven manner, offering an energy-efficient paradigm for the next generation of machine intelligence. However, the current focus within the SNN community prioritizes accuracy optimization through the development of large-scale models, limiting their viability in resource-constrained and low-power edge devices. To address this challenge, we introduce a lightweight and hardware-friendly Quantized SNN (Q-SNN) that applies quantization to both synaptic weights and membrane potentials. By significantly compressing these two key elements, the proposed Q-SNNs substantially reduce both memory usage and computational complexity. Moreover, to prevent the performance degradation caused by this compression, we present a new Weight-Spike Dual Regulation (WS-DR) method inspired by information entropy theory. Experimental evaluations on various datasets, including static and neuromorphic, demonstrate that our Q-SNNs outperform existing methods in terms of both model size and accuracy. These state-of-the-art results in efficiency and efficacy suggest that the proposed method can significantly improve edge intelligent computing. Wenjie Wei, Ammar Belatreche, Yichen Xiao, Honglin Cao, Zhenbang Ren, Guoqing Wang 0003, Malu Zhang, Yang Yang 0002 |
ACM Multimedia | 4 |
| 2024 | Spike-based Neuromorphic Model for Sound Source LocalizationabstractBiological systems possess remarkable sound source localization (SSL) capabilities that are critical for survival in complex environments. This ability arises from the collaboration between the auditory periphery, which encodes sound as precisely timed spikes, and the auditory cortex, which performs spike-based computations. Inspired by these biological mechanisms, we propose a novel neuromorphic SSL framework that integrates spike-based neural encoding and computation. The framework employs Resonate-and-Fire (RF) neurons with a phase-locking coding (RF-PLC) method to achieve energy-efficient audio processing. The RF-PLC method leverages the resonance properties of RF neurons to efficiently convert audio signals to time-frequency representation and encode interaural time difference (ITD) cues into discriminative spike patterns. In addition, biological adaptations like frequency band selectivity and short-term memory effectively filter out many environmental noises, enhancing SSL capabilities in real-world settings. Inspired by these adaptations, we propose a spike-driven multi-auditory attention (MAA) module that significantly improves both the accuracy and robustness of the proposed SSL framework. Extensive experimentation demonstrates that our SSL framework achieves state-of-the-art accuracy in SSL tasks. Furthermore, it shows exceptional noise robustness and maintains high accuracy even at very low signal-to-noise ratios. By mimicking biological hearing, this neuromorphic approach contributes to the development of high-performance and explainable artificial intelligence systems capable of superior performance in real-world environments. Dehao Zhang, Shuai Wang 0058, Ammar Belatreche, Wenjie Wei, Yichen Xiao, Haorui Zheng, Zijian Zhou 0005, Malu Zhang, Yang Yang 0002 |
NeurIPS | 5 |
| 2024 | SimpleCNN-UNet: An optic disc image segmentation network based on efficient small-kernel convolutionsabstractPathological myopia can lead to a series of eye diseases, including glaucoma and retinal pathologies. One of its most significant changes is the alteration in the size of the optic disc area in fundus images. Therefore, precise segmentation of the optic disc area is particularly important in ocular medical diagnosis. Although many well-established methods in medical image segmentation rely on Fully Convolutional Networks (FCNs), they often struggle to capture global context compared to Transformer models. However, incorporating Transformers generally necessitates larger training datasets, which can pose a significant challenge. To address these issues, Convolutional Neural Networks (CNNs) with large convolutional kernels have been proposed as an alternative for capturing contextual information, but they come with increased parameter counts and higher computational costs during training. In this paper, we introduce SimpleCNN-UNet, a lightweight image segmentation network based on small-kernel convolutions. By strategically stacking these small convolutions, we emulate the receptive field of large-kernel convolutions while substantially reducing the number of parameters. Another novel feature of SimpleCNN-UNet is the Multi-Layer Cross-Attention Gate, designed for efficient feature fusion across different levels. To overcome the limited availability of fundus image data, we employed extensive data augmentation techniques on our existing dataset. Our experimental results on the iChallenge-PM, iChallenge-AMD, iChallenge-GON, and IDRiD datasets demonstrate that SimpleCNN-UNet outperforms other image segmentation networks in terms of performance while also offering faster inference speeds and lower training costs. Yichen Xiao, Yanze Yu, Shengtao Liu, Wuzhida Bao, Shiping Wen 0001, Xingtao Zhou |
Expert Syst. Appl. | 1 |
| 2022 | Sequential Vaccine Allocation with Delayed FeedbackabstractIn this work we consider the problem of how to best allocate a limited supply of vaccines in the aftermath of an infectious disease outbreak by viewing the problem as a sequential game between a learner and an environment (specifically, a bandit problem). The difficulty of this problem lies in the fact that the payoff of vaccination cannot be directly observed, making it difficult to compare the relative effectiveness of vaccination on different population groups. Currently used vaccination policies make recommendations based on mathematical modelling and ethical considerations. These policies are static, and do not adapt as conditions change. Our aim is to design and evaluate an algorithm which can make use of routine surveillance data to dynamically adjust its recommendation. We evaluate the performance of our approach by applying it to a simulated epidemic of a disease based on real-world COVID-19 data, and show that our vaccination policy was able to perform better than existing vaccine allocation policies. In particular, we show that with our allocation method, we can reduce the number of required vaccination by at least 50% in order to keep the peak number of hospitalised patients below a certain threshold. Also, when the same batch sizes are used, our method can reduce the peak number of hospitalisation by up to 20%. We also demonstrate that our vaccine allocation does not vary the number of batches per group much, making it socially more acceptable (as it reduces uncertainty, hence results in better and more interpretable communication). Yichen Xiao, Han-Ching Ou, Haipeng Chen 0001, Van Thieu Nguyen, Long Tran-Thanh |
IJCAI | 1 |