EDBT 2026 Demo / reviewers in the wild / expert
Hu He 0001
dblp:48/4498-1
· DBLP profile ↗
18ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingabstractMemory systems characterized by high bandwidth and/or capacity alongside high access latency are becoming increasingly critical.This trend can be observed both at the device level-for instance, in non-volatile memory-and at the system level, as seen in CXL-based memory pooling architectures.To benefit from such memory in general-purpose computing systems, it is essential to employ techniques that can tolerate high memory access latency.Although prefetching has long been recognized as a classical approach for latency tolerance, conventional prefetching techniques are typically either optimized for area efficiency or constrained by limited prefetching patterns.Consequently, they often fail to convert the abundant metadata into significant performance improvements at minimal cost.To address these challenges, we propose RICH-a prefetcher that strategically consumes memory capacity and bandwidth to reduce memory access latency.First, RICH is capable of leveraging abundant metadata to improve performance by integrating spatial prefetching with diverse region sizes and prefetch triggers.Second, RICH implements such metadata with minimal overheads by employing a hierarchical on-chip/off-chip storage mechanism, thereby avoiding both large on-chip storage and critical off-chip accesses.We propose a specific implementation of RICH and evaluate it across a wide range of workloads.With increased memory latency, RICH achieves performance improvements of 8.3% over Bingo and 6.2% over PMP.This highlights the RICH's suitability for future memory systems.In a conventional system, RICH still outperforms Bingo by 3.4%. Ningzhi Ai, Wenjian He, Hu He 0001, Heng Liao, Guowei Zhang 0002 |
MICRO | 3 |
| 2025 | RISC-V-Based GPGPU With Vector Capabilities for High-Performance ComputingabstractGeneral-purpose graphics processing units (GPGPUs) have become a leading platform for accelerating modern compute-intensive applications, such as large language models and generative artificial intelligence (AI). However, the lack of advanced open-source GPGPU microarchitectures has hindered high-performance research in this area. In this article, we present Ventus, a high-performance open-source GPGPU implementation built upon the RISC-V architecture with vector extension [RISC-V vector (RVV)]. Ventus introduces customized instructions and a comprehensive software toolchain to optimize performance. We deployed the design on a field programmable gate array (FPGA) platform consisting of 4 Xilinx VU19P devices, scaling up to 16 streaming multiprocessors (SMs) and supporting 256 warps. Experimental results demonstrate that Ventus exhibits key performance features comparable to commercial GPGPUs, achieving an average of 83.9% instruction reduction and 87.4% cycle per instruction (CPI) improvement over the leading open-source alternatives. Under 4-, 8-, and 16-thread configurations, Ventus maintains robust instruction per cycle (IPC) performance with values of 0.47, 0.40, and 0.32, respectively. In addition, the tensor core of Ventus attains an extra average reduction of 69.1% in instruction count and a 68.4% cycle reduction ratio when running AI-related workloads. These findings highlight Ventus as a promising solution for future high-performance GPGPU research and development, offering a robust open-source alternative to proprietary solutions. Ventus can be found onhttps://github.com/THU-DSP-LAB/ventus-gpgpu Jingzhou Li, Fangfei Yu, Mingyuan Ma, Hualin Wu, Hu He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | HCRF: A Hardware Checkpoint-based Recovery Framework in light dual-core lockstep processorsabstractLockstep processors have become enormously popular in the global market, as automotive vehicles show significant demands for microprocessors with high reliability and safety. Vendors usually develop off-core level lockstep processors which are composed of two identical microprocessors and compare their memory access-related signals to detect transient faults. However, as safety-critical applications require harsher timeliness, the detection and recovery of this solution may not be suitable. Enlightened by the checkpoint technique in out-of-order processors to realize precise exceptions, we propose a Hardware Checkpoint-based Recovery Framework (HCRF) to shorten the detection and recovery time of light, off-core lockstep processors. The recovery time of HCRF is 41.0% faster than the baseline off-core lockstep processors and HCRF could save 39.7% of the total execution energy when transient faults occur. Jingzhou Li, Hu He 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Ventus: A High-performance Open-source GPGPU Based on RISC-V and Its Vector ExtensionabstractGeneral-purpose Graphics Processing Unit (G PG PU) has become the most popular platform for accelerating modern applications such as Large Language Models and Generative AI, while the lack of advanced open-source hardware micro architectures restricts the high-performance GPGPU research. In this work, we propose Ventus, a high-performance open-source GPGPU based on RISC- V with Vector Extension (RVV). Customized instructions and a holistic software toolchain are implemented to achieve high performance. Ventus is successfully deployed on an FPGA platform consisting of 4 Xilinx VU19P, scaling up to 16 Streaming Multiprocessors (SMs) with 256 warps. Results imply that Ventus possesses critical features of commercial GPGPUs and has achieved an average reduction of 83.9% in instruction count and 87.4% in CPI over the state-of-the-art open-source implementation. Ventus can be found on Github (https://github.com/THU-DSP-LAB/ventus-gpgpu). Jingzhou Li, Kexiang Yang, Chufeng Jin, Zexia Yang, Fangfei Yu, Mingyuan Ma, Hualin Wu, Hu He 0001 |
ICCD | 12 |
| 2023 | CLEAR: a full-stack chip-in-loop emulator for analog RRAM based computing-in-memory system
Ruihua Yu, Bin Gao 0006, Yiwen Geng, Yuyi Liu, Qingtian Zhang, Jianshi Tang, Hu He 0001, Ning Deng 0008, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 10 |
| 2023 | An Error-Free 64KB ReRAM-Based nvSRAM Integrated to a Microcontroller Unit Supporting Real-Time Program Storage and RestorationabstractNonvolatile SRAM (nvSRAM), which integrates the nonvolatile elements with SRAM using a direct bit-to-bit connection has raised much attention in the past few years, owing to its fast parallel data transfer and fast power-on/off speed. However, few nvSRAM macros have been silicon verified to be enacted through the power-failure event. On the other hand, the capacity of fabricated nvSRAM macro is small (~ Kbit) to date, inhibiting its practical application. This study presents a novel ReRAM-based nvSRAM bitcell with improved reliability and scalability. A 64KB nvSRAM macro was designed and integrated into a 32-bit microcontroller unit (MCU). The chip was fabricated using HfOx-based BEOL ReRAM and a 130nm CMOS technology. To pursue fast storage, a write-without-verify scheme is adopted to program ReRAM, measurement results show that the raw bit error rate between the power outages is < 0.1% for the full macro under such constraint. Cryptography and machine learning applications are successfully performed on the MCU system. For the first time, with the help of correction techniques, we achieved an error-free nvSRAM macro that is reliable enough to store/restore programs and demonstrated a real-time robotic control system empowered by the nvSRAM. The proposed nvSRAM macro has the largest capacity to date. Hanwen Gong, Hu He 0001, Liyang Pan, Bin Gao 0006, Jianshi Tang, Sining Pan, Dabin Wu, He Qian, Huaqiang Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | A memory neural system built based on spiking neural networkabstractMemory’s mechanism has always been the most tempting treasure for researchers. Many contributions have been delivered to unearth the mystery of memory. In this paper, we present our effort at attempting to reveal the mechanism of memory through computational neuroscience approach. We have constructed a structural efficient memory neural system with three modules, which could simulate the process where new memory is generated and kept and could be extracted. We have proved that new connections grow during the memory forming phase are vital for the keeping of memory. We propose that neurons in the memory layer could be divided into two kinds of neurons: neurons serve as interfaces for memory, and neurons serve as the main body for the keeping of memory. We also provide a method to regulate the memory layer to avoid epileptic states and work properly. The result shows our method could generate memory neural system with reasonably high memory extraction accuracy, high energy efficiency, and high robustness for different input stimulations. Hu He 0001, Xu Yang 0003, Yunlin Lei, Ning Deng 0008 |
Neurocomputing | 1 |
| 2019 | Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm AnalysisabstractRRAM based neural-processing-unit (NPU) is emerging for processing general purpose machine intelligence algorithms with ultra-high energy efficiency, while the imperfections of the analog devices and cross-point arrays make the practical application more complicated. In order to improve accuracy and robustness of the NPU, device-circuit-algorithm codesign with consideration of underlying device and array characteristics should outperform the optimization of individual device or algorithm. In this work, we provide a joint device-circuit-algorithm analysis and propose the corresponding design guidelines. Key innovations include: 1) An end-to-end simulator for RRAM NPU is developed with an integrated framework from device to algorithm. 2) The complete design of circuit and architecture for RRAM NPU is provided to make the analysis much close to the real prototype. 3) A large-scale neural network as well as other general-purpose networks are processed for the study of device-circuit interaction. 4) Accuracy loss from non-idealities of RRAM, such as I-V nonlinearity, noises of analog resistance levels, voltage-drop for interconnect, ADC/DAC precision, are evaluated for the NPU design. Xiaochen Peng, Huaqiang Wu, Bin Gao 0006, Hu He 0001, Youhui Zhang, Shimeng Yu, He Qian |
DAC | 5 |
| 2019 | On-Chip Analog Trojan Detection Framework for Microprocessor TrustworthinessabstractWith the globalization of semiconductor industry, hardware security issues have been gaining increasing attention. Among all hardware security threats, the insertion of hardware Trojans is one of the main concerns. Meanwhile, many current Trojan detection solutions follow the assumption that the hardware Trojan itself should be composed of digital logic. This assumption is invalidated by recently proposed analog Trojans which are extremely small and can detect rare events. This paper proposes a runtime hardware Trojan detection method which is geared toward detecting such advanced Trojans. The principle of this method is to guard a set of concerned signals, and initiate a hardware interrupt request when abnormal toggling events occur in these guarded signals. To prove the effectiveness of this method, we design a processor based on ARMv7-A&R ISA, and insert an analog Trojan into the processor. We fabricated the design in an SMIC 130-nm process and demonstrate the effectiveness of the proposed methodology. Yumin Hou, Hu He 0001, Kaveh Shamsi, Yier Jin, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | On Improving Performance and Energy Efficiency for Register-File Connected Clustered VLIW Architectures for Embedded System UsageabstractTraditionally, the register allocation phase for clustered very long instruction words (VLIW) architecture is implemented independently from instruction scheduling and cluster assignment. However, independently performing register allocation, scheduling and cluster assignment could have negative effect on the other phases. The research of this paper is focused on register-file connected clustered VLIW (RFCC VLIW) architecture. In RFCC VLIW architecture, the number of access ports to the global register file from each cluster and the number of registers in the global register file are both limited, due to the consideration of limited chip area of embedded systems. Thus, the distribution of inter-cluster data transferring must be carefully arranged; otherwise, there will be conflicts, which harm the performance and energy consumption. This paper proposes an algorithm to take register pressure into consideration while performing instruction scheduling and cluster assignment, so as to optimize performance and energy for embedded systems with RFCC VLIW architecture. The result shows that our algorithm can significantly reduce the penalty of performance and energy consumption due to register pressure of global register file. Hu He 0001, Xu Yang 0003 |
Comput. J. | 1 |
| 2014 | An Implementation of Message-Passing Interface over VxWorks for Real-Time Embedded Multi-Core SystemsabstractMessage-passing interface (MPI) has proved to be very successful in the high performance computing domain. However, suitability of MPI for embedded real-time system design is still under investigation. In this work, we have provided our methods and experiences of implementing MPI parallel environment for a real-time embedded multi-core system. Our main contributions were to: (1) enable hyper transport bus communication mechanism to establish MPI parallel environment; (2) support VxWorks operating system for establishing MPI parallel environment and (3) enhance the real-time property of MPI mechanism. The digital signal processor (DSP)-MPI presented in this work can also be used on other platforms supporting VxWorks operating system. The results indicate that the real-time property of DSP-MPI has been improved significantly compared with MPICH2. The test on realistic applications also shows that DSP-MPI can fulfill the requirement of our target multi-core platform. Xu Yang 0003, Deyuan Guo, Hu He 0001, Haijing Tang |
Comput. J. | 3 |
| 2014 | A Fast Application-Based Supply Voltage Optimization Method for Dual Voltage FPGAabstractDual supply voltage was a mature method to reduce the dynamic power of specific and programmable circuits, and the unsettled low voltage level (VL) was proved to have impact on its effect. In this paper, a circuit-level power model is developed to estimate the optimal VLfast for field-programmable gate array (FPGA). The model is mainly based on the path delay distribution of applications and the delay function of the integrated circuit technology. It can also count minor factors, such as path overlap, transition density, and capacitance. Experiment was conducted on a 90-nm FPGA model using MCNC benchmark. The results showed that the proposed method could generate near optimum VLfor most benchmarks. The best power reduction ratio is only 5.6% less than the gate-level heuristic method, which is relatively precise, but our method is ~100-10000 times faster. It implies that the dual voltage design with variable VL is a possible and promising low power method for field-programmable devices. Jianfeng Zhu 0001, Liyang Pan, Yaru Yan, Hu He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2011 | A chip-level path-delay-distribution based Dual-VDD method for low power FPGA (abstract only)abstractDual-VDD FPGA architecture has been proposed to reduce the FPGA's power consumption, where a low VDD (VDDL) is assigned to non-critical resources and unused resources are power-gated. In this paper, a path-delay-distribution (PDD) based design method of supply voltage in dual-VDD FPGA is developed, which gives an estimated optimal VDD solution for the required applications. Meanwhile, an improved tree-based VDD assignment algorithm is accordingly designed. Thus chip-level optimization of dual-VDD FPGA is achieved on the chosen granularity with the power consumption minimized. Based on MCNC benchmark circuits at 90nm technology node, our experimental result shows that: the power reduction rate depends on VDDL level; the design method proposed in this work gives the optimal one automatically. This design method could be utilized to guide the FPGA automatic design, saving the time to search for the system's optimal supply voltage, and the proposed assignment algorithm is more efficient in dynamic power reduction. Jianfeng Zhu 0001, Yaru Yan, Hu He 0001, Liyang Pan |
FPGA | 5 |
| 2011 | A cost-efficient self-configurable BIST technique for testing multiplexer-based FPGA interconnect
Jianfeng Zhu 0001, Hu He 0001, Liyang Pan |
J. Electron. Test. | 2 |
| 2011 | Erratum to: A Cost-Efficient Self-Configurable BIST Technique for Testing Multiplexer-Based FPGA Interconnect
Jianfeng Zhu 0001, Hu He 0001, Liyang Pan |
J. Electron. Test. | 2 |
| 2009 | Simultaneous Multithreading VLIW DSP Architecture with Dynamic Dispatch MechanismabstractThis paper presents a novel simultaneous multithreading (SMT) VLIW DSP architecture with dynamic dispatch mechanism to address the challenge of the underutilization of computing resources in the non-unit assumed latency (NUAL) VLIW DSPs. The SMT technology exploits the unused instruction slots by converting the thread-level parallelism to the instruction-level parallelism, improving the efficiency. With the specifically designed registers for eliminating the horizontal dependencies among the execution-packet, the NUAL VLIW DSP architecture supports issuing any subset of instructions of the execution-packet based on the availability of the corresponding functional units. With the dynamic dispatch mechanism, the DSP issues instructions to functional unit at run-time rather than at compile-time, such that the issue conflicts among multiple threads are reduced significantly. The new VLIW DSP architecture is implemented and evaluated, and the results show that the architecture can effectively increase the processor throughput, hide the cache miss latencies, and improve the performance on digital signal processing. Zheng Shen, Hu He 0001, Yihe Sun |
DSD | 2 |
| 2007 | A 2-Dimension Force-Directed Scheduling Algorithm for Register-File-Connectivity Clustered VLIW ArchitectureabstractVery long instruction word (VLIW) processors are generally implemented as clustered architectures in order to reduce delay, area and power when function units increase. However, a side effect of clustered architectures is processor performance degradation, due to additional latency and copy operations from data transfers between these clusters. Therefore, appropriate scheduling algorithms must be applied to overcome this. This paper presents a new register-file-connectivity clustered VLIW (RFCC-VLIW) architecture, in which a global register file is used to transfer data between clusters. Using the global register file, latency and copy operations can be eliminated. Additionally, the paper presents a scheduling algorithm for the RFCC-VLIW architecture. The two-dimension force-directed algorithm can assign instructions evenly to all clusters and reduce usage of ports and global registers. Experimental results show our algorithm outperforms other scheduling algorithms for clustered VLIW when targeted toward RFCC-VLIW architectures. Zhixiong Zhou, Hu He 0001, Yihe Sun, Adriel Cheng |
ASAP | 2 |
| 2005 | A new register file access architecture for software pipelining in VLIW processorsabstractThis paper presents a novel architecture of register files that combines the local register files and the global register file for clustered VLIW (Very Long Instruction Word) processors. The communication between function units through global register file will be more efficient. The concept of associate register is introduced for this architecture. This makes it possible to write a result to two destination registers in one operation, which can efficiently speed up the software pipelining. Hu He 0001, Yihe Sun |
ASP-DAC | 2 |