EDBT 2026 Demo / reviewers in the wild / expert
Huaien Gao
dblp:32/5771
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0001-0528-3076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Efficiency Bidirectional Translator between SystemC and VerilogabstractThe SystemC language, with its higher level of abstraction, plays a critical role in facilitating hardware/software co-design and architecture exploration. However, as most hardware models are predominantly written in Verilog and translating between SystemC and Verilog remains a challenge, an efficient and reliable tool for translating between these two languages is essential to streamline system development. This article proposes SCAV, a bidirectional translator between SystemC and Verilog, which breaks these limitations. SCAV provides a fully automated solution for translating both SystemC to Verilog and Verilog to SystemC, leveraging a translation framework with front-end/back-end separation. Additionally, SCAV incorporates an Abstract Syntax Tree (AST) filter, optimizing the translation process by filtering out invalid content. The experimental results demonstrate that SCAV achieves a 100% adaptation rate for Verilog and a 98% adaptation rate for SystemC, with 100% accuracy in both directions. Furthermore, SCAV outperforms existing tools, delivering a minimum speedup of 18% across various test cases. Xin Zheng 0001, Yongfeng Zhong, Shaofen Zeng, Huaien Gao, Shuting Cai, Xiaoming Xiong |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2026 | FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V CoresabstractArea efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available athttps://github.com/LX-IC/FPUltra Xian Lin, Jiahao Lan, Xin Zheng 0001, Huanxin Zhuang, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Efficient Design Space Exploration for the BOOM Using SAC-Based Reinforcement LearningabstractDesign space exploration (DSE) is crucial for optimizing the performance, power, and area (PPA) of CPU microarchitectures ($\mu $-archs). While various machine learning (ML) algorithms have been applied to the$\mu $-arch DSE problem, the potential of reinforcement learning (RL) remains underexplored. In this article, we propose a novel RL-based approach to address the reduced instruction set computer V (RISC-V) CPU$\mu $-arch DSE problem. This approach enables dynamic selection and optimization of$\mu $-arch parameters without relying on predefined modification sequences, thus significantly enhancing exploration flexibility. To address the challenges posed by high-dimensional action spaces and sparse rewards, we use a discrete soft actor-critic (SAC) framework with entropy maximization to promote efficient exploration. In addition, we integrate multistep temporal-difference (TD) learning, an experience replay (ER) buffer, and return normalization to improve sample efficiency and learning stability during training. Our method further aligns optimization with user-defined preferences by normalizing PPA metrics relative to baseline designs. Experimental results on the Berkeley out-of-order machine (BOOM) demonstrate that the proposed approach achieves superior performance compared with state-of-the-art methods, showcasing its effectiveness and efficiency for$\mu $-arch DSE. Our code is available athttps://github.com/exhaust-create/SAC-DSE. Mingjun Cheng, Xin Zheng 0001, Xian Lin, Huaien Gao, Shuting Cai, Xiaoming Xiong, Bei Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | FPUx: High-Performance Floating-Point Support for Cost-Constrained RISC-V CoresabstractIn the Internet of Things (IoT) field, cloud and fog computing dramatically increase the complexity of floating-point (FP) calculations. Cost-constrained microcontrollers (MCUs) urgently need more efficient FP computing methods, such as integrated FP units (FPUs). To this end, this brief proposes FPUx, a high-performance FPU designed through a hybrid pipeline and state-machine approach. The FPUx is integrated into E203 for implementation (E203-FPUx). Furthermore, the Easy-lite is proposed to reduce handshake delay and a range of single-precision FP (FP32) arithmetic IPs are designed to customize FPUs. Compared with E203-FPnew and E203, the performance of E203-FPUx is improved by$1.5\times $and$36\times $, and the total energy consumption is saved by 36% and 1430% on average, respectively. Xian Lin, Heming Liu, Xin Zheng 0001, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | BSSE: Design Space Exploration on the BOOM With Semi-Supervised LearningabstractWith the rising prominence of RISC-V-based microprocessors in processor design, the challenge of exploring the vast and complex RISC-V microarchitecture design space has become increasingly apparent. We propose the Berkeley Out-of-Order Machine Semi-Supervised Explorer (BSSE)—a novel framework leveraging the semi-supervised learning method and parallel emulation to speed up and make tradeoffs on the RISC-V microarchitecture design space exploration (DSE). BSSE constructs the initial training dataset with the microarchitecture experimental design sampling (MEDS) method and then employs the cotraining-style k-nearest neighbors (Co-KNN) model to fit the microarchitecture features to the architectural metric value space. The trained Co-KNN model assists in searching a Pareto-optimal set with parallel emulation. Finally, a distance-based method is proposed to select a designer-preferred microarchitecture from the identified Pareto-optimal set. Extensive experiments on the Berkeley Out-of-Order Machine (BOOM) show that our proposed BSSE method can search for a better Pareto-optimal set with less time consumption compared to the state-of-the-art methods and can find microarchitectures that are equivalent to or even better than the existing manually designed BOOM microarchitectures. Xin Zheng 0001, Mingjun Cheng, Jiasong Chen, Huaien Gao, Xiaoming Xiong, Shuting Cai |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Conditional prediction of time series using spiral recurrent neural network
Huaien Gao, Rudolf Sollacher |
ESANN | 1 |
| 2008 | Efficient online learning with Spiral Recurrent Neural NetworksabstractDistributed intelligent systems like self-organizing wireless sensor and actuator networks are supposed to work mostly autonomous even under changing environmental conditions. This requires robust and efficient self-learning capabilities implementable on embedded systems with limited memory and computational power. We present a new solution called spiral recurrent neural networks with an online learning based on an extended Kalman filter and gradients as in real-time recurrent learning. We illustrate its performance using artificial and real-life time series and compare it to other approaches. Finally we describe a few potential applications. Rudolf Sollacher, Huaien Gao |
IJCNN | 2 |
| 2007 | Spiral Recurrent Neural Network for Online Learning
Huaien Gao, Rudolf Sollacher, Hans-Peter Kriegel |
ESANN | 1 |