Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Huwan Peng

dblp:172/2773 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-5855-2228ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Virtual and augmented reality · 50% Rendering · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 67% Reconfigurable computing and FPGAs · 33%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
collaborative rendering
0.512021
Q-VR: system-level design for future mobile collaborative virtual reality · ASPLOS 2021
Virtual and augmented reality › virtual reality
mobile virtual reality
0.512021
Q-VR: system-level design for future mobile collaborative virtual reality · ASPLOS 2021
Compilers and program optimization
deep learning compiler
0.312018
Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.312018
Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.312018
Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312018
Exploring the programmability for deep learning processors: from architecture to tensorization · DAC 2018

Methods — techniques the papers use, named apart from their topics

software-hardware co-design · 1.5dynamic collaborative rendering · 1.5
YearPublicationVenuePosition
2025 ReaLLM: A Trace-Driven Framework for Rapid Simulation of Large-Scale LLM Inference
abstract
As Large Language Models (LLMs) continue to scale, optimizing their deployment requires efficient hardware and system co-design. However, current LLM performance evaluation frameworks fail to capture both chip-level execution details and system-wide behavior, making it difficult to assess realistic performance bottlenecks. In this work, we introduce ReaLLM, a trace-driven simulation framework designed to bridge the gap between detailed accelerator design and large-scale inference evaluation. Unlike prior simulators, ReaLLM integrates kernel profiling derived from detailed microarchitectural simulations with a new trace-driven end-to-end system simulator, enabling precise evaluation of parallelism strategies, batching techniques, and scheduling policies. To address the high computational cost of exhaustive simulations, ReaLLM constructs a precomputed kernel library based on hypothesized scenarios, interpolating results to efficiently explore a vast design space of LLM inference systems. Our validation against real hardware demonstrates the framework's accuracy, achieving an average end-to-end latency prediction error of only 9.1% when simulating inference tasks running on 4 NVIDIA H100 GPUs. We further use ReaLLM to evaluate popular LLMs' end-to-end performance across traces from different applications and identify key system bottlenecks, showing that modern GPU-based LLM inference is increasingly compute-bound rather than memory-bandwidth bound at large scale. Additionally, we significantly reduce simulation time with our precomputed kernel library by a factor of$6 \times$for full-simulations and$164 \times$for workload SLO exploration. ReaLLM is open-source and available at https://github.com/bespoke-silicon-group/reallm.
Huwan Peng, Scott Davidson 0004, Chuanjin Richard Shi, Michael B. Taylor
ASAP1
2021 Q-VR: system-level design for future mobile collaborative virtual reality
abstract
High Quality Mobile Virtual Reality (VR) is what the incoming graphics technology era demands: users around the world, regardless of their hardware and network conditions, can all enjoy the immersive virtual experience. However, the state-of-the-art software-based mobile VR designs cannot fully satisfy the realtime performance requirements due to the highly interactive nature of user's actions and complex environmental constraints during VR execution. Inspired by the unique human visual system effects and the strong correlation between VR motion features and realtime hardware-level information, we propose Q-VR, a novel dynamic collaborative rendering solution via software-hardware co-design for enabling future low-latency high-quality mobile VR. At software-level, Q-VR provides flexible high-level tuning interface to reduce network latency while maintaining user perception. At hardware-level, Q-VR accommodates a wide spectrum of hardware and network conditions across users by effectively leveraging the computing capability of the increasingly powerful VR hardware. Extensive evaluation on real-world games demonstrates that Q-VR can achieve an average end-to-end performance speedup of 3.4x (up to 6.7x) over the traditional local rendering design in commercial VR devices, and a 4.1x frame rate improvement over the state-of-the-art static collaborative rendering.
Chenhao Xie 0001, Xie Li, Yang Hu 0001, Huwan Peng, Michael B. Taylor, Shuaiwen Song
ASPLOS4
2018 Exploring the programmability for deep learning processors: from architecture to tensorization
abstract
This paper presents an instruction and Fabric Programmable Neuron Array (iFPNA) architecture, its 28nm CMOS chip prototype, and a compiler for the acceleration of a variety of deep learning neural networks (DNNs) including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected (FC) networks on chip. The iFPNA architecture combines instruction-level programmability as in an Instruction Set Architecture (ISA) with logic-level reconfigurability as in a Field-Programmable Gate Array (FPGA) in a sliced structure for scalability. Four data flow models, namely weight stationary, input stationary, row stationary and tunnel stationary, are described as the abstraction of various DNN data and computational dependence. The iFPNA compiler partitions a large-size DNN to smaller networks, each being mapped to, optimized and code generated for, the underlying iFPNA processor using one or a mixture of the four data-flow models. Experimental results have shown that state-of-art large-size CNNs, RNNs, and FC networks can be mapped to the iFPNA processor achieving the near ASIC performance.
Chixiao Chen, Huwan Peng, Xindi Liu, Chuanjin Richard Shi
DAC2