EDBT 2026 Demo / reviewers in the wild / expert
Xuliang Yu
dblp:271/9589
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-8414-6960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MambaOPU: An FPGA Overlay Processor for State-space-duality-based Mamba ModelsabstractState-space models (SSMs), such as Mamba, have emerged as a promising alternative to Transformers. However, the recently developed Mamba2, based on state space duality (SSD), is highly memorybound and suffers from limited computation efficiency. This inefficiency arises from its irregular broadcast element-wise multiplications and structured sparse computations. In this work, we propose MambaOPU, an FPGA overlay processor, to accelerate SSD. First, to reduce memory overhead, we introduce a software-hardware co-optimized operator fusion framework. Specifically, operator merging combines adjacent broadcast multiplication and summation operations into a single descriptor, while operator backward shifting embeds segment multiplication into subsequent operations. Both techniques shorten the computation path and improve computation efficiency. Second, to enhance sparse computation efficiency, we skip zero-region computations using a tensor-reorder-and-group algorithm combined with a sparse-predefined data fetcher. Additionally, since Mamba integrates linear operations with SSD, we develop a reconfigurable systolic array to improve data reuse across different computation modes. Extensive experiment results demonstrate that MambaOPU achieves up to $1812 \times$ and $880.79 \times$ higher normalized throughput and up to $12908 \times$ and $24.27 \times$ higher energy efficiency over Intel Xeon Gold 6348 CPU and NVIDIA A100 GPU, respectively. Shaoqiang Lu, Xuliang Yu, Tiandong Zhao, Siyuan Miao, Xinsong Sheng, Ting-Jung Lin, Lei He 0001 |
DAC | 2 |
| 2025 | FeKAN: Efficient Kolmogorov-Arnold Networks Accelerator Using FeFET-based CAM and LUTabstractKolmogorov-Arnold networks (KANs) have emerged as a promising alternative to MLP due to their adaptive learning capabilities for complex dependencies through B-spline basis activations (BBA). However, existing in-memory accelerators optimized for MLP-based DNNs are primarily designed for vector-matrix multiplication (VMM), making them inefficient for the dynamic and recursive B-spline interpolation (BSI) operations required by KANs. In this work, we propose FeKAN, an FeFET-based architecture designed to accelerate BBA operations. First, we develop a software-hardware co-optimized framework for mapping B-spline basis functions (BBF), leveraging a two-stage design space exploration (DSE) algorithm in combination with FeFET-based Look-Up Tables (LUT) and Content-Addressable Memory (CAM). This framework translated dynamic BSI operations into static codebook lookups, achieving a balanced trade-off between memory and computational efficiency. Second, we propose compress-sparsity-column (CSC) based encoding for B-spline basis function and grouped-computation strategy for memory and energy reduction. Third, we propose a groupedpipeline optimization strategy to mitigate data dependencies, significantly enhancing computation efficiency. Experimental results demonstrate that FeKAN achieves up to $150.68 \mathrm{~K} \times$ and $4664 \times$ higher throughput and up to $606.87 \times$ and $11196 \times$ greater energy efficiency over Intel Xeon Silver 4310 CPU and NVIDIA A6000 GPU, respectively. Xuliang Yu, Yu Qian 0002, Xunzhao Yin, Cheng Zhuo, Liang Zhao 0004 |
DAC | 1 |
| 2025 | Probabilistic Person-in-Bed Detection Using Accelerometer SignalsabstractUsing accelerometer data in smart bed systems offers a cost-effective solution for person-in-bed detection. In this work, we propose a lightweight probabilistic model for this task. The accelerometer time series is first divided into multiple patches, with high-frequency noise filtered through a combined optimization of 1D convolution and spectral pooling. An LSTM-based feature extraction module is then employed to capture temporal dependencies. Subsequently, a Fourier classification head is applied to generate probabilistic detection outputs. The proposed architecture achieves an accuracy of 1.0 on the segmented detection task and 0.915 on the streaming detection task in the ICASSP 2025 signal Processing Grand Challenge, organized by the Analog Garage. The implementation is available at https://github.com/JiayiGao04/person-in-bed-detection. Kaite Shi, Jiayi Gao, Xuliang Yu |
ICASSP | 4 |
| 2024 | AESHA: Accelerating Eigen-decomposition-based Sparse Transformer with Hybrid RRAM-SRAM ArchitectureabstractCompute-in-memory (CIM) architectures based on emerging nonvolatile memories (eNVM) are recognized as promising candidates for the efficient computation of self-attention-based Transformer models, which are bounded by memory-intensive dynamic computations involving large matrices. However, existing CIM-based approaches mostly focused on the acceleration of vector-matrix multiplications (VMM) to obtain Q, K and V matrices for direct or sparse calculations of the vanilla Transformer. By storing large amounts of intermediate results in eNVM, the endurance limit is ignored which contradicts the reality of these technologies. In this work, a hybrid RRAM-SRAM CIM architecture to accelerate Transformers, namely AESHA, is proposed based on the eigen-decomposition of WQWTK and the systolic in-memory attention reconstruction. By utilizing the features of symmetric/skew-symmetric real matrices, AESHA translates dynamic attention computation to RRAM-based sparse feature transformation and SRAM-based systolic attention reconstruction to fully exploit the advantages of both on the architectural level. Experiments on a broad spectrum of benchmarks demonstrate that AESHA delivers superior performance in terms of pipeline optimization and energy efficiency. Specifically, AESHA achieves 2.59×, 2.69×, 2.62×, 2.58×, 3.08× and 7.84× maximum static RRAM memory footprint reduction in BERT-Base, BERT-Large, BigBird, Sanger, ViT-Base and ViT-Large over vanilla computation stack, respectively. It eliminates all runtime write access to the RRAM arrays to circumvent the endurance limit. Additionally, AESHA demonstrates 3170.0×, 95.6×, 4.1×, 19.0×, and 19.2× speedup, and 336.0K×, 12.6K×, 7.3×, 23.1×, and 25.8× energy reduction over CPU, GPU, ReBERT, ReTransformer, and CPSAA, respectively. Xuliang Yu, Tianwei Ni, Xinsong Sheng, Lei He 0001 |
ICCAD | 1 |
| 2024 | An SRAM Compute-In-Memory Macro Based on Direct Coupling SAR ADC and DAC ReuseabstractDeep neural networks (DNNs) can be implemented in low-power, small-area systems using compute-in-memory (CIM). This study proposes a CIM MAC macro based on 8T SRAM and a 5-bit Successive Analog-to-Digital Data Converter (SAR ADC) readout circuit to execute low-power and highspeed calculations. This architecture has no buffers between the low-power SAR ADC and the pre-stage. The SAR ADC’s capacitors are reused throughout calculating operations. The proposed macro is implemented using 55nm technology and achieves an INT4 throughput of 23.27 GOPS per ADC and an energy efficiency of 49.85 TOPS/W with 4-bit input, 4-bit weight, and 5-bit output. Yongteng Ma, Xuliang Yu, Zhichao Tan |
ISCAS | 2 |
| 2023 | Sample Intercorrelation-Based Multidomain Fusion Network for Aquatic Human Activity Recognition Using Millimeter-Wave RadarabstractTo address the issue that existing multi-domain fusion methods do not consider data correlation within mini-batch data, the first attempt is made in this letter to propose a method based on sample inter-correlation learning multi-domain fusion network (SIMFNet), which aims to accommodate multivariate domain data and further enhance the aquatic human activity recognition performance. To fully utilize the radar multidimensional information, the three-branch convolution neural network (CNN) feature extractor is first employed to extract domain-specific features from the time-range map (TRM), time-Doppler map (TDM) and cadence velocity diagram (CVD). Then, the multi-domain features are fused and fed into the graph construction layer (GCL) to generate instance graphs. Next, a graph aggregation layer (GAL) is applied to aggregate node information from various-hop neighborhood domains. Finally, node-level classification is used to achieve aquatic human activity recognition. The experimental results evaluated on the built aquatic human activity recognition dataset demonstrate that the proposed SIMFNet has better generalization performance than the state-of-the-art multi-domain fusion methods. Xuliang Yu, Zhihui Cao, Zhijing Wu 0003, Chunyi Song, Zhiwei Xu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | People Flow Detection Algorithm Based on a Multiencoder-Classifier Cotraining Architecture for FMCW RadarabstractDeep learning (DL) frameworks are widely used in various applications due to their superiority over conventional handcrafted feature-based algorithms. However, applying DL to time-range feature map-based people flow detection (PFD) with frequency modulated continuous wave (FMCW) radar still faces several challenges: 1) simultaneously achieving people counting and motion direction recognition requires a unified framework, 2) existing mainstream network backbones designed for semantic information-rich optical images or natural language suffer performance loss in time-range feature maps with weak semantic information, and 3) the construction of labeled data sets in PFD scenes is usually costly, and limited data lead to performance loss due to overfitting. Therefore, this paper proposes novel solutions from various aspects to efficiently apply DL to time-range feature map-based PFD. First, new preprocessing pipelines with Doppler spectrum analysis-based feature map truncation are proposed for the first time to simultaneously achieve people counting and direction recognition using a single radar in the radar PFD field. Second, a novel lightweight multiscale feature space fusion-based convolutional neural network backbone (MFSNet) is designed to efficiently extract multichannel differentiated representative features from time-range feature maps. Finally, a multiencoder-classifier cotraining architecture based on embedded features with two data synthesis methods and a newly designed loss function is proposed to improve the generalization ability of the algorithm. Using the test data set collected from real scenes, the performance comparison results show that the proposed PFD algorithm outperforms the state-of-the-art algorithms and ablation studies demonstrate the effectiveness of each component of the proposed algorithm in PFD. Zhihui Cao, Zhijing Wu 0003, Xuliang Yu, Chunyi Song, Zhiwei Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Novel Potential Drowning Detection System Based on Millimeter-Wave RadarabstractRadar is widely used in human activity recognition because of its powerful micro-doppler feature capture capability and environmental adaptability. In this work, we propose a novel radar-based potential drowning detection system. To enhance the cross-domain fusion efficiency and intra-domain feature learning, we design a two-stage fusion network for the drowning detection system. In the first-stage fusion, we integrate the encoded features of three-domain radar maps along either the temporal or spatial dimension. In the second-stage fusion, we use Attention-LSTM and 1D-CNN to extract deep information from temporal-fused and spatial-fused features, and further combine these features using a trainable weighted average strategy. Based on our proposed novel fusion architecture, fine-grained aquatic human activity recognition is achieved. In the experiments, we collect a nine-class aquatic human activity dataset. The experimental results demonstrate the superiority of the proposed TSFNet over the state-of-the-art models. The dataset and the associated codes are available at: https://github.com/DingdongD/aquatic-activity-dataset. Xuliang Yu, Zhihui Cao, Zhijing Wu 0003, Chunyi Song, Jiang Zhu 0004, Zhiwei Xu 0003 |
ICARCV | 1 |