EDBT 2026 Demo / reviewers in the wild / expert
Yuhan She
dblp:329/9496
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-3748-577XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chariot: Compiler-Aware Heterogeneous Graph Representation Learning for Automated HLS OptimizationabstractHigh-level synthesis (HLS) design space exploration (DSE) aims to find Pareto-optimal designs but is hindered by slow synthesis evaluations. Existing graph neural network (GNN) surrogates struggle with homogeneous-style graph representations (causing signal over-squashing) and imprecise source-level heuristics for pragma mapping. We propose Chariot, an automated HLS optimization framework. Chariot leverages LLVM-based static analysis for high-fidelity Use-Def chain tracking, modeling HLS designs as semantic-rich heterogeneous graphs that explicitly map directives to true hardware targets. Our framework achieves state-of-the-art QoR prediction, identifying Pareto-optimal solutions with drastically reduced ranking regret while delivering orders-of-magnitude DSE speedup. Jierui Liu, Yuhan She, Rongliang Fu, Tsung-Yi Ho, Hong Yan 0001, Ray C. C. Cheung |
FCCM | 3 |
| 2026 | ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGAabstractVision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs), yet efficiently deploying them on FPGAs remains challenging. Linear layers struggle with dynamic activation outliers that render static quantization ineffective, while uniform quantization fails to capture the weight distribution at low bit-widths. Furthermore, while associative scan accelerates SSMs on GPUs, its memory access patterns are misaligned with the streaming dataflow required by FPGAs. To address these challenges, we present ViM-Q1, a scalable algorithm-hardware co-design for end-to-end ViM inference on the edge. We introduce a hardware-aware quantization scheme combining dynamic per-token activation quantization and per-channel smoothing to mitigate outliers, alongside a custom 4-bit per-block Additive Power-of-Two (APoT) weight quantization. The models are deployed on a runtime-parameterizable FPGA accelerator featuring a linear engine employing a Lookup-Table (LUT) unit to replace multiplications with shift-add operations, and a fine-grained pipelined SSM engine that parallelizes the state dimension while preserving sequential recurrence. Crucially, the hardware supports runtime configuration, adapting to diverse dimensions and input resolutions across the ViM family. Implemented on an AMD ZCU102 FPGA, ViM-Q achieves an average 4.96× speedup and 59.8× energy efficiency gain over a quantized NVIDIA RTX 3090 GPU baseline for low-batch inference on ViM-tiny. This co-design shows a viable path for deploying ViM models on resource-constrained edge devices. Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu |
FCCM | 2 |
| 2026 | ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba ModelsabstractState-space models (SSMs), such as Mamba, provide an efficient alternative to Transformers for vision tasks by replacing their quadratic-cost self-attention with linear complexity state update. However, efficiently deploying Vision Mamba (ViM) models on FPGA platforms is challenging, as the latency is dominated by two key components: linear layers and the selective SSM. For the linear layers, highly dynamic activation outliers across tokens render conventional static quantization techniques ineffective. Meanwhile, while the associative scan algorithm is effective in accelerating SSM on GPUs, its data access pattern is fundamentally mismatched with FPGA architectures when mapping the model's inherently sequential recurrence, creating a critical dataflow bottleneck. Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu |
FPGA | 2 |
| 2026 | SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel EstimationabstractChannel estimation is crucial in 5G communication networks for optimizing transmission parameters and ensuring reliable, high-speed communication. However, the use of multiple-input and multiple-output (MIMO) and millimeter-wave (mmWave) in 5G networks presents challenges in achieving accurate estimation under strict latency requirements on resource-limited hardware platforms. To address these challenges, we proposeSwiftChannel, an algorithm-hardware co-design framework that integrates a hardware-friendly deep learning-based channel estimator with a dedicated accelerator. Our approach employs a convolutional neural network enhanced with a parameter-free attention mechanism, which effectively reconstructs full-resolution spatial-frequency domain channel matrices from low-resolution least squares (LS) estimates. We further develop a multi-stage model compression pipeline combining knowledge distillation, convolution re-parameterization, and quantization-aware training, resulting in substantial model size reduction with negligible accuracy loss. The hardware accelerator, implementing the compressed model and the LS estimator on FPGA platforms using High-level Synthesis (HLS), features a fine-grained pipeline architecture and optimized dataflow strategies. Tested on a Zynq UltraScale+ RFSoC, the accelerator achieves sub-millisecond latency, providing up to 24x speed-up and over 33x improvement in energy efficiency compared to GPU-based solutions. Extensive evaluations demonstrate that the proposed design generalizes not only across various noise levels and user mobilities, but also to a variety of unseen channel profiles, outperforming state-of-the-art baselines. By unifying algorithmic innovation with hardware-aware design, our work presents a future-proof channel estimation solution for 5G MIMO systems. The source codes for the dataset synthesis, deep learning algorithm, and HLS-based FPGA design are accessible via GitHub. Shengzhe Lyu, Yuhan She, Di Duan, Tao Ni 0003, Yu Hin Chan, Chengwen Luo 0001, Ray C. C. Cheung, Weitao Xu |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | CAHLS: Source-to-Source Transformation to Generate Cycle Accurate Models for High-Level SynthesisabstractHigh-Level Synthesis (HLS) empowers the ability to synthesize a customized hardware description from an untimed software description. However, the quality of the generated hardware is affected by the HLS tool. Current state-of-the-art commercial HLS tools adopt static-scheduling-based algorithms, which perform well for the regular designs but suffer performance degradation for the control-dominant designs. Dynamic scheduling, on the other hand, performs well for control flows but loses certain optimizations, like resource sharing and critical path optimizations, resulting in area overhead and frequency drop. In this paper, we propose a source-to-source transformation to generate an equivalent pseudo cycle-accurate model, so that 1) the transformed code runs dynamically based on different control conditions, and 2) the transformed code still fits in the static HLS tool. As future work, this transformation can be integrated into a compiler to automatically optimize the control-dominant designs in the static-scheduling HLS flow. Yuhan She, Jierui Liu, Ray C. C. Cheung, Hong Yan 0001 |
CODES+ISSS | 1 |
| 2025 | A Speculative Loop Pipeline Framework with Accurate Path Modeling for High-Level SynthesisabstractLoop pipelining is a key optimization in high-level synthesis (HLS), aimed at overlapping the execution of iterations. Static scheduling, dominant in commercial HLS tools, configures the pipeline based on compile-time analysis, proving conservative for designs with irregular control flow and memory access due to imbalanced recurrences. Speculative Loop pipeline (SLP) is a novel concept that addresses the problem by introducing the speculation and recovery mechanism at the source level to improve the throughput. Although proven promising, it has a significant gap from practical application: It requires accurate early-stage modeling of the pipeline configuration for each path, which is unable to obtain with classic HLS scheduling methods because the SLP process itself interferes with the path length. In this work, we made a step forward by proposing a practical SLP framework with accurate path modeling ability through iterative tuning. We further optimize the SLP technology by combining automatic dataflow extraction with speculative source-level transformation to further boost the performance in specific design patterns. Our framework works on the source level and is easy to be plugged into existing downstream HLS tools. Experiment results demonstrate significant performance improvements over commercial HLS tools and better resource trade-offs compared to the state-of-the-art dynamic-scheduling-based solutions. Yuhan She, Jierui Liu, Ray C. C. Cheung, Hong Yan 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Melting Glacier: A 37-Year (1984-2020) High-Resolution Glacier-Cover Record of MT. KilimanjaroabstractCommonly recognized as an important symbol of the tropics and global warming, the glacier loss on Mt. Kilimanjaro has received worldwide attention for decades. In this paper, we propose a high-resolution glacier-cover (GC) record of Mt. Kilimanjaro over the period from 1984 to 2020, using a novel deep learning-based semantic segmentation method and Google Earth images, as well as digital elevation model (DEM) and ERA5-Land (ERA5) for snowline and temperature variations analysis. Our method achieves an accuracy of 94.37%, which proves the model's capability to record the GC areas precisely. The results show that (1) the GC area dramatically decreases from 19.2 km2to 3.6 km2during 37 years, which decreases about 4% and 2% per year from 1984 to 2000 and from 2000 to 2020 respectively, (2) the snowline altitude rises from$4,651 m$to$5,088 m$by about$437 m$, and (3) the average$5,000 m$air temperature on Mt. Kilimanjaro increases from −2.1 °C to −1.1 °C by about 1 °C. This study indicates that there will be no GC within a few decades if the current loss continues. Shuai Yuan 0005, Juepeng Zheng, Lixian Zhang 0002, Runmin Dong, Yile Xing, Yuhan She, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 6 |