EDBT 2026 Demo / reviewers in the wild / expert
Shangkun Li
dblp:319/7718
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAPO: Design Structure-Aware Pass Ordering for HLS via Contrastive and Reinforcement LearningabstractHigh-Level Synthesis (HLS) tools are widely adopted in FPGA-based domain-specific accelerator design. However, existing tools rely on fixed optimization strategies inherited from software compilations, limiting their effectiveness. Tailoring optimization strategies to specific designs requires deep semantic understanding, accurate hardware metric estimation, and advanced search algorithms - capabilities that current approaches lack.We propose DAPO, a design structure-aware pass ordering framework that extracts program semantics from control and data flow graphs, employs contrastive learning to generate rich embeddings, and leverages an analytical model for accurate hardware metric estimation. These components jointly guide a reinforcement learning agent to discover design-specific optimization strategies. Evaluations on standard HLS benchmarks demonstrate that our end-to-end flow delivers 1.67× speedup on pragma-free designs and a 2.36× speedup on designs with pragmas over Vitis HLS with comparable resource usage. Jinming Ge, Linfeng Du, Likith Anaparty, Shangkun Li, Tingyuan Liang, Afzal Ahmad, Vivek Chaturvedi, Sharad Sinha, Zhiyao Xie, Jiang Xu 0001, Wei Zhang 0012 |
DATE | 4 |
| 2026 | A Cluster-Based Distributed Memory Architecture for CGRAs
Shangkun Li, Cheng Tan 0002, Jinming Ge, Linfeng Du, Jiang Xu 0001, Wei Zhang 0012 |
DATE | 1 |
| 2026 | NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are a promising and versatile accelerator platform, offering a balance between the performance and efficiency of specialized accelerators and software programmability. However, their full potential is severely hindered by control flow in accelerated kernels, as control flow (e.g., loops, branches) is fundamentally incompatible with the parallel, data-driven CGRA fabric. Prior strategies to resolve this mismatch in CGRA kernel acceleration are either inefficient, sacrificing performance for generality, or lack generality due to the difficulty of adapting them across different execution models. Thus, a general and unified solution for efficient CGRA kernel acceleration remains elusive. This paper introduces NEURA, a unified and retargetable compilation framework that systematically resolves the control-dataflow mismatch in CGRAs. NEURA's core innovation is a novel, pure dataflow intermediate representation (IR) built on a predicated type system. In this IR, control contexts are embedded as a predicate within each data, making control an intrinsic property of data. This mechanism enables NEURA to systematically flatten complex control flow into a single unified dataflow graph. This unified representation decouples kernel representation from hardware, empowering NEURA to retarget diverse CGRAs with different execution models and microarchitectural features. When targeted to a high-performance spatio-temporal CGRA, NEURA delivers a 2.20x speedup on kernel benchmarks and up to 2.71x geometric mean speedup on real-world applications over state-of-the-art (SOTA) high-performance baselines. It also provides a competitive solution against the SOTA low-power CGRA when retargeted to a spatial-only CGRA. NEURA is open-source and available at https://github.com/coredac/neura. Shangkun Li, Jinming Ge, Diyuan Tao, Linfeng Du, Jiang Xu 0001, Wei Zhang 0012, Cheng Tan 0002 |
Proc. ACM Program. Lang. | 1 |
| 2026 | ESFA: An Efficient Scalable FFT Design Framework on Versal AI EngineabstractThe fast fourier transform (FFT) is widely used to convert a time-domain signal into its frequency-domain representation in various fields. Previous works have demonstrated efficient FFT implementation on various accelerators. The emergence of AI Engines (AIE) on AMD Xilinx’s Versal ACAP brings the possibility of further improvement in computing efficiency. However, previous solutions have been restricted to a single-AIE manner, which limits the FFT size and neglects the potential of employing multiple AIEs. This paper proposes the ESFA framework, which can efficiently and automatically implement a scalable FFT on the Versal ACAP with multiple AIEs. The framework includes an analytical model to report the quality of results (QoRs) estimation for legal FFT partition modes, comprehensively covering the throughput-resource trade-off choices across the design space. In addition, the layout problem is formulated in an ILP to enhance the area efficiency. The framework also incorporates an automatic code generator to enable an agile implementation of the desired design. Our experiments on the VCK190 board show that we achieve 9,226/2,059MS/ssimulation/system throughput on the 1K-point FFT with a data width of 32, which obtains up to 10.1x speedup compared with AMD Xilinx’s library targeting AIE, meanwhile, 17.5x, 23.2x, and 0.9x speedup compared to the state-of-the-art designs on ASIC, CGRA, FPGA. Linfeng Du, Shangkun Li, Wei Zhang 0012 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Robin: RWKV Accelerator using Block Circulant Matrices based on FPGAabstractRecent advancements in linear-attention models, such as RWKV, have opened up new possibilities for efficient sequence processing by reducing the computational overhead of traditional Transformer architectures. Field-programmable gate arrays (FPGAs) offer a compelling solution for deep learning applications by providing customizable hardware architectures that enhance computational efficiency and flexibility. However, deploying these models on FPGAs introduces several challenges. Previous FPGA deployment workflows tend to focus on general machine learning tasks, lacking sufficient integration between software and hardware optimizations. Besides, FPGAs are constrained by limited on-chip and off-chip memory, posing significant challenges for weight storage. Moreover, the predominance of linear operations on the GPU runtime leads to significant computational bottlenecks. These obstacles necessitate innovative solutions to bridge the performance gap between FPGAs and GPUs while preserving model accuracyTo overcome these challenges, we introduce Robin, a fine-grained FPGA accelerator workflow that integrates both algorithm-level and hardware-level optimization. Robin leverages a weight compression technique based on Partial Block Circulant Matrices (PBCM), which effectively reduces storage demands while maintaining accuracy. Based on PBCM, our design employs a configurable circulant computing core that fully exploits the bit-width efficiency of DSP48E resources through two DSP packaging strategies to support both circulant and standard matrix operations. The combined end-to-end software-hardware co-design enables Robin to achieve up to a 3.09× increase in throughput and a 7.31× boost in energy efficiency compared to high-end Tesla A100 GPU implementations, making it a compelling solution for deploying RWKV models on FPGAs. Shangkun Li, Chuyi Dai, Chaofang Ma, Wei Zhanh |
ICCAD | 2 |
| 2025 | A Three Level Interleaved DC-DC Converter for EV Battery Charging from a Residential PV SystemabstractIn a domestic photovoltaic (PV)-battery system, the battery stack serves as a backup power source, supplying household loads overnight and storing the excess solar energy but its expensive due to battery cost. On the other hand, Electric Vehicles which are becoming more popular, have an embedded battery that could be used as power buffers if these could be conveniently interfaced to the residential PV system.This paper proposes a bidirectional three-level interleaved (B3LI) DC-DC converter that can facilitate interconnection of the EV battery to the DC link of a PV inverter and avoids the need for an expensive bidirectional front end AC/DC charger. The main advantages of this topology are the utilization of multilevel and interleaving techniques which reduce magnetics size and the phase shedding control, which reduces the switching losses at low powers, which may be characteristic for the way the EV’s battery would operate when connected to the household PV system. The proposed B3LI converter is experimentally evaluated in a lab setting relevant to the applications. Shangkun Li, Perdana Putera, Christian Klumpner, Mark Sumner, Mohamed Hajj |
IECON | 1 |
| 2024 | FADO: Floorplan-Aware Directive Optimization Based on Synthesis and Analytical Models for High-Level Synthesis Designs on Multi-Die FPGAsabstractMulti-die FPGAs are widely adopted for large-scale accelerators, but optimizing high-level synthesis designs on these FPGAs faces two challenges. First, the delay caused by die-crossing nets creates an NP-hard floorplanning problem. Second, traditional directive optimization cannot consider resource constraints on each die or the timing issue incurred by the die-crossings. Furthermore, the high algorithmic complexity and the large scale lead to extended runtime for legalizing the floorplan of HLS designs under different directive configurations. To co-optimize the directives and floorplan of HLS designs on multi-die FPGAs, we formulate the co-search based on bin-packing variants and present two iterative optimization flows. The first (FADO 1.0) relies on a pre-built QoR library. It involves a greedy, latency-bottleneck-guided directive search, and an incremental floorplan legalization. Compared with a global floorplanning solution, it takes 693X~4925X shorter search time and achieves 1.16X~8.78X better design performance, measured in workload execution time. To remove the time-consuming QoR library generation, the second flow (FADO 2.0) integrates an analytical QoR model and redesigns the directive search to accelerate convergence. Through experiments on mixed dataflow and non-dataflow designs, compared with 1.0, FADO 2.0 further yields a 1.40X better design performance on average after implementation on the Alveo U250 FPGA. Linfeng Du, Tingyuan Liang, Jinming Ge, Shangkun Li, Sharad Sinha, Jieru Zhao, Zhiyao Xie, Wei Zhang 0012 |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2021 | The Investigation of Behavior Change in EEG Signals During Induction of AnesthesiaabstractAnesthesiology aims to make anesthesia safer and increase the precision of prognoses. Correct assessment of the anesthesia depth is crucial to its safety. At present, intraoperative electroencephalogram (EEG) monitoring is the primary mode of anesthesia depth monitoring and judgment. However, most clinical anesthesiologists rely on commercial anesthesia depth monitors to judge anesthesia depth, such as bispectral index (BIS) and patient state index (PSI). This may lack an understanding of associated changes in brain wave quantization. Therefore, this study conducts quantitative analyses of EEG signals during anesthesia induction. EEG signals are processed within specific time windows and extracted brainpower density spectrum arrays with different frequency bands, brain electrical signal spectra, source frequencies and other key indicators. Analysis and comparison of these indicators clarifies patterns of variation in EEG signals during early anesthesia induction. The spectral edge frequencies (SEFs) of EEG signals within different time windows can be modeled accurately, from which the specific time points of EEG signal changes are derived. Furthermore, the relationship between patient age and the effect of anesthetic drugs is preliminarily investigated by analyzing the SEF variations of different age groups. This study quantifies changes in the EEG signals of patients at the initial stage of anesthesia induction and drug-related effects are observed, which opens a way for further exploration of EEG changes in patients under general anesthesia. Shangkun Li, Wei Dan, Lihao Chen, Su Min |
Int. J. Pattern Recognit. Artif. Intell. | 1 |