EDBT 2026 Demo / reviewers in the wild / expert
Feng Yu 0003
dblp:28/1708-3
· DBLP profile ↗
17ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-9740-2537ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
Feng Yu 0003, Hongshi Tan, Yao Chen 0008, Weng-Fai Wong, Bingsheng He |
ISCA | 1 |
| 2026 | Weakly supervised single-channel interictal epileptiform discharge detection with spatial priors
Jiale Shen, Feng Yu 0003 |
Expert Syst. Appl. | 5 |
| 2026 | Learning to Detect Sleep Micro-Events From Coarse Sleep Stage AnnotationsabstractSleep micro-events, such as sleep spindles and K-complexes, are closely associated with neurological cognitive functions. While artificial intelligence (AI)-assisted sleep micro-event detection provides automated annotation to reduce reliance on labor-intensive expert labeling, current supervised approaches require precisely annotated datasets that remain scarce in clinical practice. To overcome this data bottleneck, this paper introduces a Weakly Supervised Sleep Micro-Event Detector (WSSMED) that leverages readily available coarse sleep stage annotations. The proposed WSSMED features a dual-branch architecture, consisting of a wave prototype module and a cluster module, designed to capture the fine-grained sleep micro-event patterns experts rely on for sleep staging. This framework infers expert logic from coarse annotations while mitigating performance degradation caused by annotation inconsistencies arising from inter-rater variability. Experiments conducted on two public datasets and one clinical dataset demonstrate that WSSMED achieves state-of-the-art performance in detecting sleep spindles and K-complexes, as evaluated at both sample-level and event-level in terms of precision, recall and F1-score metrics. Furthermore, subject-level evaluation demonstrates that the density and duration of micro-events detected by WSSMED-key metrics linked to cognitive function and neurological status-align more closely with expert annotations than those of other reported methods. These results highlight the clinical potential of WSSMED for reliable sleep micro-event analysis. Yan Pei 0004, Chengyang Han, Lisan Zhang, Feng Yu 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Configurable DSP-Based CAM Architecture for Data-Intensive Applications on FPGAsabstractContent-addressable memory (CAM) is a type of fast memory unique in its ability to perform parallel searches of stored data based on content rather than specific memory addresses. They have been used in many domains, such as networking, databases, and graph processing. Field-programmable gate arrays (FPGAs) are an attractive platform for implementing CAMs because of their low latency, reconfigurability, and energy-efficient nature. However, such implementations also face significant challenges, including high resource utilization, limited scalability, and suboptimal performance due to the extensive use of look-up tables (LUTs) and block RAMs (BRAMs). These issues stem from the inherent limitations of FPGA architectures when handling the parallel operations required by CAMs, often leading to inefficient designs that cannot meet the demands of high-speed, data-intensive applications. To address these challenges, we propose a novel configurable CAM architecture that leverages the digital signal processing (DSP) blocks available in modern FPGAs as the core resource. By utilizing DSP blocks’ data storage and logic capabilities, our approach enables configurable CAM architecture with efficient multi-query support while significantly reducing search and update latency for data-intensive applications. The DSP-based CAM architecture offers enhanced scalability, higher operating frequency, and improved performance compared to traditional LUT and BRAM-based designs. In addition, we demonstrate the effectiveness of our proposed CAM architecture with a triangle counting application on real graphs. This innovative use of DSP blocks also opens up new possibilities for highperformance, data-intensive applications on FPGAs. Our proposed design is open-sourced at: https://github.com/Xtra-Computing/DSP_CAM/. Yao Chen 0008, Feng Yu 0003, Weng-Fai Wong, Bingsheng He |
DAC | 2 |
| 2025 | Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph ProcessingabstractEfficient graph processing is critical in various modern applications, such as social network analysis, recommendation systems, and large-scale data mining. Traditional single-FPGA systems struggle to handle the increasing size and complexity of real-world graphs due to limitations in memory and computational resources. Existing multi-FPGA solutions face significant challenges, including high communication overhead caused by irregular data transfer patterns and workload imbalances stemming from skewed graph distributions. These inefficiencies hinder scalability and performance, highlighting a critical research gap. To address these issues, we introduce Clementi, an efficient multi-FPGA graph processing framework that features customized fine-grained pipelines for computation and cross-FPGA communication. Clementi uniquely integrates an accurate performance model for execution time prediction, enabling a novel scheduling method that balances workload distribution and minimizes communication overhead by overlapping communication and computation stages. Experimental results demonstrate that Clementi achieves speedups of up to 8.75× compared to state-of-the-art multi-FPGA designs, indicating significant improvements in processing efficiency as the number of FPGAs increases. This near-linear scalability underscores the framework' s potential to enhance graph processing capabilities in practical applications. Clementi is open-sourced at https://github.com/Xtra-Computing/Clementi. Feng Yu 0003, Hongshi Tan, Xinyu Chen 0001, Yao Chen 0008, Bingsheng He, Weng-Fai Wong |
Proc. ACM Manag. Data | 1 |
| 2025 | WaveSleepNet: An Interpretable Network for Expert-Like Sleep StagingabstractAlthough deep learning algorithms have proven their efficiency in automatic sleep staging, their "black-box" nature has limited their clinical adoption. In this study, we propose WaveSleepNet, an interpretable neural network for sleep staging that reasons in a similar way to sleep clinical experts. In this network, we utilize the latent space representations generated during training to identify characteristic wave prototypes corresponding to different sleep stages. The feature representation of an input signal is segmented into patches within the latent space, each of which is compared against the learned wave prototypes. The proximity between these patches and the wave prototypes is quantified through scores, indicating the prototypes' presence and relative proportion within the signal. The scores serve as the decision-making criteria for final sleep staging. During training, an ensemble of loss functions is employed for the prototypes' diversity and robustness. Furthermore, the learned wave prototypes are visualized by analyzing occlusion sensitivity. The efficacy of WaveSleepNet is validated across three public datasets, achieving sleep staging performance that are on par with those of the state-of-the-art models. A detailed case study examining the decision-making process of WaveSleepNet demonstrates that it aligns closely with American Academy of Sleep Medicine (AASM) manual guidelines. Another case study systematically explained the misidentified reasons behind each sleep stage. WaveSleepNet's transparent process provides specialists with direct access to the physiological significance of the model's criteria, allowing for future validation, adoption and further enrichment by sleep clinical experts. Yan Pei 0004, Feng Yu 0003, Lisan Zhang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Area-Efficient Pipeline Architecture for Serial Real-Valued Fast Fourier TransformabstractThis brief presents a novel pipeline architecture designed to compute the fast Fourier transform (FFT) on real input signals in a serial format. This architecture significantly improves resource efficiency by sharing adders between butterfly and rotator structures. In addition, a novel data management approach for N-point radix-2 serial real-valued FFT (RFFT) has been proposed, which not only simplifies the data reordering circuit between processing elements (PEs) but also achieves natural order data output. The real-valued 1024-point FFT has been implemented on a field-programmable gate array (FPGA). Compared with typical real-valued serial commutator (RSC) FFT architecture, the proposed architecture achieves substantial improvement, including a reduction of 10.3% in the number of lookup tables (LUTs) and 12.5% in flip-flops (FFs). Kun Li 0032, Hongji Fang, Zhen-guo Ma, Feng Yu 0003, Bo Zhang 0097, Qianjian Xing |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A Fast Floating-Point Multiply-Accumulator Optimized for Sparse Linear Algebra on FPGAsabstractThis brief presents a pipelined floating-point Multiply–Accumulator (FPMAC) architecture designed to accelerate sparse linear algebra operations. By designing a lookup-table-based 5–3 carry-save adder (CSA) and combining it with a 3–2 CSA, the proposed design minimizes the critical path and boosts operational speed. Moreover, the proposed architecture takes advantage of data characteristics in sparse linear algebra to displace the shift unit in the critical accumulation loop, further increasing the throughput rate. In addition, the integration of a lookup-table-based leading-zero anticipator (LZA) enhances normalization efficiency. Experimental results show that, compared with reported FPMAC designs, the proposed architecture may achieve a significantly higher maximum clock frequency for single-precision floating-point operations. Kun Li 0032, Xiangyu Hao, Zhen-guo Ma, Feng Yu 0003, Bo Zhang 0097, Qianjian Xing |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar DisordersabstractPrevious research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requirements of clinical diagnosis. Efficient multimodal fusion strategies have great potential for applications in multimodal data and can further improve the performance of medical diagnosis models. In this work, we utilize both sMRI and fMRI data and propose a novel multimodal diagnosis model for bipolar disorder. The proposed Patch Pyramid Feature Extraction Module extracts sMRI features, and the spatio-temporal pyramid structure extracts the fMRI features. Finally, they are fused by a fusion module to output diagnosis results with a classifier. Extensive experiments show that our proposed method outperforms others in balanced accuracy from 0.657 to 0.732 on the OpenfMRI dataset, and achieves the state of the art. Sheng Shi, Shan An, Fengmei Fan, Wenshu Ge, Feng Yu 0003, Zhiren Wang |
ICASSP | 7 |
| 2024 | DTP-Net: Learning to Reconstruct EEG Signals in Time-Frequency Domain by Multi-Scale Feature ReuseabstractElectroencephalography (EEG) signals are prone to contamination by noise, such as ocular and muscle artifacts. Minimizing these artifacts is crucial for EEG-based downstream applications like disease diagnosis and brain-computer interface (BCI). This paper presents a new EEG denoising model, DTP-Net. It is a fully convolutional neural network comprising Densely-connected Temporal Pyramids (DTPs) placed between two learnable time-frequency transformations. In the time-frequency domain, DTPs facilitate efficient propagation of multi-scale features extracted from EEG signals of any length, leading to effective noise reduction. Comprehensive experiments on two public semi-simulated datasets demonstrate that the proposed DTP-Net consistently outperforms existing state-of-the-art methods on metrics including relative root mean square error (RRMSE) and signal-to-noise ratio improvement ( ∆SNR). Moreover, the proposed DTP-Net is applied to a BCI classification task, yielding an improvement of up to 5.55% in accuracy. This confirms the potential of DTP-Net for applications in the fields of EEG-based neuroscience and neuro-engineering. An in-depth analysis further illustrates the representation learning behavior of each module in DTP-Net, demonstrating its robustness and reliability. Yan Pei 0004, Qianhao Chen, Feng Yu 0003, Lisan Zhang |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Learning longer-term dependencies via grouped distributor unit
Feng Yu 0003 |
Neurocomputing | 2 |
| 2015 | An Optimum Architecture for Continuous-Flow Parallel Bit ReversalabstractWith the aim of minimizing memory and latency, this letter presents a novel bit-reversal architecture for continuous-flow parallel pipelined FFT processors. It harnesses the theory that any permutation can be decomposed to a series of elementary bit-exchanges. The main contribution of this letter are twofold. First, it achieves continuous-flow bit reversal in parallel with the minimum memory and minimum latency. Second, the architecture, composed of memory and 2-to-1 multiplexers, are simple and regular for general power-of-2 parallelism. Furthermore, it supports different common radices, including radix-2,radix-4, and radix-8. Feng Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Block Region of Interest Method for Real-Time Implementation of Large and Scalable Image ReconstructionabstractIn field programmable gate array (FPGA)-based reconstruction of high-resolution (several-gigabyte) images, architectures that store the entire source image in block memory are no longer feasible. In this letter, a block region of interest (ROI) method is proposed for large and scalable image reconstruction, using 2D 4 × 4 truncated sinc interpolation by storing only a slice of the source image in block memory. To ensure a high hit rate in block memory for interpolation and improve memory access efficiency, this method provides an effective criterion for partitioning, which subdivides the output image into blocks based on the ROI in the source image. Moreover, a special storage pattern is designed to simultaneously provide 16 pixels for each interpolation to enable real-time implementation by full pipelining with no memory data redundancy. The proposed method was validated by experimental results on an FPGA board along with the corresponding simulations on MATLAB. Various image sizes were tested to prove its flexibility. The time required to reconstruct an output image of 32 k × 32 k at 200 MHz was 2.79 s, thereby achieving a processing speed of 385 Mpixels/s in single floating point precision. Feng Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2015 | A Combined SDC-SDF Architecture for Normal I/O Pipelined Radix-2 FFTabstractWe present an efficient combined single-path delay commutator-feedback (SDC-SDF) radix-2 pipelined fast Fourier transform architecture, which includes log2N - 1 SDC stages, and 1 SDF stage. The SDC processing engine is proposed to achieve 100% hardware resource utilization by sharing the common arithmetic resource in the time-multiplexed approach, including both adders and multipliers. Thus, the required number of complex multipliers is reduced to log4N - 0.5, compared with log2N - 1 for the other radix-2 SDC/SDF architectures. In addition, the proposed architecture requires roughly minimum number of complex adders log2N + 1 and complex delay memory 2N + 1.5log2N - 1.5. Zeke Wang, Xue Liu 0003, Bingsheng He, Feng Yu 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | A Family of Fast Hadamard-Fourier Transform AlgorithmsabstractIn this letter, we present a family of fast Hadamard-Fourier transform algorithms which combined Walsh Hadamard and discrete Fourier transforms into one single algorithm. These family algorithms can be computed in butterfly structure, and have similar sparse matrix factorization in each stage, and have less computation stages than the sum of Walsh Hadamard and discrete Fourier transforms. We factorize the algorithms with regular sparse matrices for every stage in radix-R mode, where R is power of 2. Teng Su, Feng Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2011 | A pipelined architecture for normal I/O order FFTabstractWe present a novel pipelined fast Fourier transform (FFT) architecture which is capable of producing the output sequence in normal order. A single-path delay commutator processing element (SDC PE) has been proposed for the first time. It saves a complex adder compared with the typical radix-2 butterfly unit. The new pipelined architecture can be built using the proposed processing element. The proposed architecture can lead to 100% hardware utilization and 50% reduction in the overall number of adders required in the conventional pipelined FFT designs. In order to produce the output sequence in normal order, we also present a bit reverser, which can achieve a 50% reduction in memory usage. Xue Liu 0003, Feng Yu 0003, Zeke Wang |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | An efficient radix-2 fast Fourier transform processor with ganged butterfly engines on field programmable gate arraysabstractWe present a novel method to implement the radix-2 fast Fourier transform (FFT) algorithm on field programmable gate arrays (FPGA). The FFT architecture exploits parallelism by having more pipelined units in the stages, and more parallel units within a stage. It has the noticeable advantages of high speed and more efficient resource utilization by employing four ganged butterfly engines (GBEs), and can be well matched to the placement of the resources on the FPGA. We adopt the decimation-infrequency (DIF) radix-2 FFT algorithm and implement the FFT processor on a state-of-the-art FPGA. Experimental results show that the processor can compute 1024-point complex radix-2 FFT in about 11 µs with a clock frequency of 200 MHz. Zhen-guo Ma, Feng Yu 0003, Ruifeng Ge, Zeke Wang |
J. Zhejiang Univ. Sci. C | 2 |