EDBT 2026 Demo / reviewers in the wild / expert
Hui Wang 0166
dblp:39/721-166
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0000-4078-6667ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 65% Embedded and real-time systems · 17% Reconfigurable computing and FPGAs · 10% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
2.5 | 3 | 2025 | BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models · DAC 2025 Pushing the Limits of BFP on Narrow Precision LLM Inference · AAAI 2025 MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Pushing the Limits of BFP on Narrow Precision LLM Inference · AAAI 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Pushing the Limits of BFP on Narrow Precision LLM Inference · AAAI 2025 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.9 | 1 | 2025 | Pushing the Limits of BFP on Narrow Precision LLM Inference · AAAI 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
0.9 | 1 | 2025 | Pushing the Limits of BFP on Narrow Precision LLM Inference · AAAI 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM quantization accelerator |
0.9 | 1 | 2025 | BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models · DAC 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.9 | 1 | 2025 | NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.8 | 1 | 2024 | MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024 |
Embedded and real-time systems › real-time scheduling › mixed-criticality scheduling
mixed-criticality systems |
0.8 | 1 | 2024 | MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024 |
Embedded and real-time systems
real-time scheduling |
0.8 | 1 | 2024 | MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024 |
Memory systems
cache |
0.3 | 1 | 2025 | NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025 |
Memory systems › cache
cache miss reduction |
0.3 | 1 | 2025 | NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025 |
Memory systems › cache
prefetching |
0.3 | 1 | 2025 | NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025 |
Methods — techniques the papers use, named apart from their topics
block floating point · 2.6lookup table · 1.7hardware-software co-design · 1.7speculative execution · 0.9runahead execution · 0.9quantization · 0.9timing analysis · 0.8context switching · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pushing the Limits of BFP on Narrow Precision LLM InferenceabstractThe substantial computational and memory demands of Large Language Models (LLMs) hinder their deployment. Block Floating Point (BFP) has proven effective in accelerating linear operations, a cornerstone of LLM workloads. However, as sequence lengths grow, nonlinear operations, such as Attention, increasingly become performance bottlenecks due to their quadratic computational complexity. These nonlinear operations are predominantly executed using inefficient floating-point formats, which renders the system challenging to optimize software efficiency and hardware overhead. In this paper, we delve into the limitations and potential of applying BFP to nonlinear operations. Given our findings, we introduce a hardware-software co-design framework (DB-Attn), including: (i) DBFP, an advanced BFP version, overcomes nonlinear operation challenges with a pivot-focus strategy for diverse data and an adaptive grouping strategy for flexible exponent sharing. (ii) DH-LUT, a novel lookup table algorithm dedicated to accelerating nonlinear operations with DBFP format. (iii) An RTL-level DBFP-based engine is implemented to support DB-Attn, applicable to FPGA and ASIC. Results show that DB-Attn provides significant performance improvements with negligible accuracy loss, achieving 74% GPU speedup on Softmax of LLaMA and 10x low-overhead performance improvement over SOTA designs. Hui Wang 0166, Xiaomeng Han, Zhengpeng Zhao, Zhe Jiang 0004 |
AAAI | 1 |
| 2025 | BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language ModelsabstractLarge language models (LLMs), with their billions of parameters, pose substantial challenges for deployment on edge devices, straining both memory capacity and computational resources. Block Floating Point (BFP) quantisation reduces memory and computational overhead by converting high-overhead floating point operations into low-bit fixed point operations. However, BFP requires aligning all data to the maximum exponent, which causes loss of small and moderate values, resulting in quantisation error and degradation in the accuracy of LLMs. To address this issue, we propose a Bidirectional Block Floating Point (BBFP) data format, which reduces the probability of selecting the maximum as shared exponent, thereby reducing quantisation error. By utilizing the features in BBFP, we present a full-stack Bidirectional Block Floating Point-Based Quantisation Accelerator for LLMs (BBAL), primarily comprising a processing element array based on BBFP, paired with proposed cost-effective nonlinear computation unit. Experimental results show BBAL achieves a 22% improvement in accuracy compared to an outlier-aware accelerator at similar efficiency, and a 40% efficiency improvement over a BFP-based accelerator at similar accuracy. Xiaomeng Han, Jing Wang 0113, Junyang Lu, Hui Wang 0166, X. x. Zhang, Ning Xu 0009, Zhe Jiang 0004 |
DAC | 5 |
| 2025 | NVR: Vector Runahead on NPUs for Sparse Memory AccessabstractDeep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4 x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16 KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the $\mathbf{L 2}$ cache size by the same amount. Hui Wang 0166, Zhengpeng Zhao, Jing Wang 0113, Yushu Du, Chenhao Ma 0006, Xiaomeng Han, Dean You, Jiapeng Guan, Zhe Jiang 0004 |
DAC | 1 |
| 2025 | MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded ProcessorsabstractRunahead execution is a technique to mask memory latency caused by irregular memory accesses. By pre-executing the application code during occurrences of long-latency operations and prefetching anticipated cache-missed data into the cache hierarchy, runahead effectively masks memory latency for subsequent cache misses and achieves high prefetching accuracy; however, this technique has been limited to superscalar out-of-order and superscalar in-order cores. For implementation in scalar in-order cores, the challenges of area-/energy-constraint and severe cache contention remain. Here, we build the first full-stack system featuring runahead, MERE , from SoC and a dedicated ISA to the OS and programming model. Through this deployment, we show that enabling runahead in scalar in-order cores is possible, with minimal area and power overheads, while still achieving high performance. By re-constructing the sequential runahead employing a hardware/software co-design approach, the system can be implemented on a mature processor and SoC. Building on this, an adaptive runahead mechanism is proposed to mitigate the severe cache contention in scalar in-order cores. Combining this, we provide a comprehensive solution for embedded processors managing irregular workloads. Our evaluation demonstrates that the proposed MERE attains 93.5% of a 2-wide out-of-order core’s performance while constraining area and power overheads below 5%, with the adaptive runahead mechanism delivering an additional 20.1% performance gain through mitigating the severe cache contention issues. Dean You, Jieyu Jiang, Yushu Du, Zhihang Tan, Hui Wang 0166, Jiapeng Guan, Shuai Zhao 0004, Zhe Jiang 0004 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2024 | MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSsabstractModern Mixed-Criticality Systems (MCSs) rely on hardware heterogeneity to satisfy ever-increasing computational demands. However, most of the heterogeneous co-processors are designed to achieve high throughput, with their micro-architectures executing the workloads in a streaming manner. This streaming execution is often non-preemptive or limited-preemptive, preventing tasks’ prioritisation based on their importance and resulting in frequent occurrences of algorithmic priority and/or criticality inversions. Such problems present a significant barrier to guaranteeing the systems’ real-time predictability, especially when co-processors dominate the execution of the workloads (e.g., DNNs and transformers).In contrast to existing works that typically enable coarse-grained context switch by splitting the workloads/algorithms, we demonstrate a method that provides fine-grained context switch on a widely used open-source DNN accelerator by enabling instruction-level preemption without any workloads/algorithms modifications. As a systematic solution, we build a real system, i.e., Make Each Switch Count (MESC), from the SoC and ISA to the OS kernel. A theoretical model and analysis are also provided for timing guarantees. Experimental results reveal that, compared to conventional MCSs using non-preemptive DNN accelerators, MESC achieved a 250 x and 300 x speedup in resolving algorithmic priority and criticality inversions, with less than 5% overhead. To our knowledge, this is the first work investigating algorithmic priority and criticality inversions for MCSs at the instruction level. Jiapeng Guan, Dean You, Yingquan Wang, Ruizhe Yang, Hui Wang 0166, Zhe Jiang 0004 |
RTSS | 6 |