EDBT 2026 Demo / reviewers in the wild / expert
Huwan Peng
dblp:172/2773
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-5855-2228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Virtual and augmented reality · 50% Rendering · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 67% Reconfigurable computing and FPGAs · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer networks
1 paper |
Edge and fog computing · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering
collaborative rendering |
0.5 | 1 | 2021 | Q-VR: system-level design for future mobile collaborative virtual reality · ASPLOS 2021 |
Virtual and augmented reality › virtual reality
mobile virtual reality |
0.5 | 1 | 2021 | Q-VR: system-level design for future mobile collaborative virtual reality · ASPLOS 2021 |
Compilers and program optimization
deep learning compiler |
0.3 | 1 | 2018 | Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.3 | 1 | 2018 | Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.3 | 1 | 2018 | Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.3 | 1 | 2018 | Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018 |
Methods — techniques the papers use, named apart from their topics
software-hardware co-design · 1.5dynamic collaborative rendering · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReaLLM: A Trace-Driven Framework for Rapid Simulation of Large-Scale LLM InferenceabstractAs Large Language Models (LLMs) continue to scale, optimizing their deployment requires efficient hardware and system co-design. However, current LLM performance evaluation frameworks fail to capture both chip-level execution details and system-wide behavior, making it difficult to assess realistic performance bottlenecks. In this work, we introduce ReaLLM, a trace-driven simulation framework designed to bridge the gap between detailed accelerator design and large-scale inference evaluation. Unlike prior simulators, ReaLLM integrates kernel profiling derived from detailed microarchitectural simulations with a new trace-driven end-to-end system simulator, enabling precise evaluation of parallelism strategies, batching techniques, and scheduling policies. To address the high computational cost of exhaustive simulations, ReaLLM constructs a precomputed kernel library based on hypothesized scenarios, interpolating results to efficiently explore a vast design space of LLM inference systems. Our validation against real hardware demonstrates the framework's accuracy, achieving an average end-to-end latency prediction error of only 9.1% when simulating inference tasks running on 4 NVIDIA H100 GPUs. We further use ReaLLM to evaluate popular LLMs' end-to-end performance across traces from different applications and identify key system bottlenecks, showing that modern GPU-based LLM inference is increasingly compute-bound rather than memory-bandwidth bound at large scale. Additionally, we significantly reduce simulation time with our precomputed kernel library by a factor of$6 \times$for full-simulations and$164 \times$for workload SLO exploration. ReaLLM is open-source and available at https://github.com/bespoke-silicon-group/reallm. Huwan Peng, Scott Davidson 0004, Chuanjin Richard Shi, Michael B. Taylor |
ASAP | 1 |
| 2021 | Q-VR: system-level design for future mobile collaborative virtual realityabstractHigh Quality Mobile Virtual Reality (VR) is what the incoming graphics technology era demands: users around the world, regardless of their hardware and network conditions, can all enjoy the immersive virtual experience. However, the state-of-the-art software-based mobile VR designs cannot fully satisfy the realtime performance requirements due to the highly interactive nature of user's actions and complex environmental constraints during VR execution. Inspired by the unique human visual system effects and the strong correlation between VR motion features and realtime hardware-level information, we propose Q-VR, a novel dynamic collaborative rendering solution via software-hardware co-design for enabling future low-latency high-quality mobile VR. At software-level, Q-VR provides flexible high-level tuning interface to reduce network latency while maintaining user perception. At hardware-level, Q-VR accommodates a wide spectrum of hardware and network conditions across users by effectively leveraging the computing capability of the increasingly powerful VR hardware. Extensive evaluation on real-world games demonstrates that Q-VR can achieve an average end-to-end performance speedup of 3.4x (up to 6.7x) over the traditional local rendering design in commercial VR devices, and a 4.1x frame rate improvement over the state-of-the-art static collaborative rendering. Chenhao Xie 0001, Xie Li, Yang Hu 0001, Huwan Peng, Michael B. Taylor, Shuaiwen Song |
ASPLOS | 4 |
| 2018 | Exploring the programmability for deep learning processors: from architecture to tensorizationabstractThis paper presents an instruction and Fabric Programmable Neuron Array (iFPNA) architecture, its 28nm CMOS chip prototype, and a compiler for the acceleration of a variety of deep learning neural networks (DNNs) including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected (FC) networks on chip. The iFPNA architecture combines instruction-level programmability as in an Instruction Set Architecture (ISA) with logic-level reconfigurability as in a Field-Programmable Gate Array (FPGA) in a sliced structure for scalability. Four data flow models, namely weight stationary, input stationary, row stationary and tunnel stationary, are described as the abstraction of various DNN data and computational dependence. The iFPNA compiler partitions a large-size DNN to smaller networks, each being mapped to, optimized and code generated for, the underlying iFPNA processor using one or a mixture of the four data-flow models. Experimental results have shown that state-of-art large-size CNNs, RNNs, and FC networks can be mapped to the iFPNA processor achieving the near ASIC performance. Chixiao Chen, Huwan Peng, Xindi Liu, Chuanjin Richard Shi |
DAC | 2 |