EDBT 2026 Demo / reviewers in the wild / expert
Pengfei Gou
dblp:96/9639
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Timestamp as a prior: Enhancing long-term time series forecast via temporal semantic-aligned contrastive learning
Pengfei Gou, Yifei Tang, Xuyun Nie, Mengjuan Liu |
Expert Syst. Appl. | 1 |
| 2023 | Orinoco: Ordered Issue and Unordered Commit with Non-Collapsible QueuesabstractModern out-of-order processors call for more aggressive scheduling techniques such as priority scheduling and out-of-order commit to make use of increasing core resources. Since these approaches prioritize the issue or commit of certain instructions, they face the conundrum of providing the capacity efficiency of scheduling structures while preserving the ideal ordering of instructions. Traditional collapsible queues are too expensive for today's processors, while state-of-the-art queue designs compromise with the pseudo-ordering of instructions, leading to performance degradation as well as other limitations. Dibei Chen, Tairan Zhang, Yi Huang 0036, Jianfeng Zhu 0001, Yang Liu 0326, Pengfei Gou, Chunyang Feng, Shaojun Wei, Leibo Liu |
ISCA | 6 |
| 2023 | MapZero: Mapping for Coarse-grained Reconfigurable Architectures with Reinforcement Learning and Monte-Carlo Tree SearchabstractCoarse-grained reconfigurable architecture (CGRA) has become a promising candidate for data-intensive computing due to its flexibility and high energy efficiency. CGRA compilers map data flow graphs (DFGs) extracted from applications onto CGRAs, playing a fundamental role in fully exploiting hardware resources for acceleration. Yet the existing compilers are time-demanding and cannot guarantee optimal results due to the traversal search of enormous search spaces brought about by the spatio-temporal flexibility of CGRA structures and the complexity of DFGs. Inspired by the amazing progress in reinforcement learning (RL) and Monte-Carlo tree search (MCTS) for real-world problems, we consider constructing a compiler that can learn from past experiences and comprehensively understand the target DFG and CGRA. Yi Huang 0036, Jianfeng Zhu 0001, Xingchen Man, Yang Liu 0326, Chunyang Feng, Pengfei Gou, Minggui Tang, Shaojun Wei, Leibo Liu |
ISCA | 7 |
| 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) can provide both energy efficiency and performance for embedded systems, and thus they are increasingly deployed in the areas of aerospace, automotive engineering, and security where reliability is also a main criterion. However, the state-of-the-art fault-tolerant strategies for CGRAs apply either temporal or spatial scheme, including redundancy, periodic detection, workload balancing, and reconfiguration, failing to exploit the feature of dynamic and partial reconfiguration of CGRAs. Also, vulnerable judging circuits and inflexible mode shifting bottleneck the reliability design of fault-tolerant CGRAs. This article proposes a novel multimode fault-tolerant framework for CGRAs, which combines spatial-redundant data paths with temporal-redundant voters and thus reduces the vulnerable judging circuits while balancing the performance and reliability. This framework can also enable a changing reliability level at runtime via an online configuration transformation method based on precompiled patterns. Within the proposed framework, we systematically searched the design space spanning various combinations of the mainstream schemes with a Markov process model to compare the effectiveness and accordingly selected five points as available modes in our design after comprehensive consideration of fault tolerance and time overhead on CGRA. The framework is comprehensively evaluated on a cycle-accurate CGRA simulator, considering both permanent and transient faults. The experimental results show that the fault coverage rate of single transient faults or permanent faults has increased from 71.74% to 93.84%, which means the fault tolerance of the system has been increased by 31.03% compared with the state-of-the-art methods. There is also a great improvement in mean-time-to-failure (MTTF) and reconfiguration latency over baseline designs. Jianfeng Zhu 0001, Xingchen Man, Guihuan Song, Yi Huang 0036, Chenchen Deng, Pengfei Gou, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | An Energy-Efficient Approximate Divider Based on Logarithmic Conversion and Piecewise Constant ApproximationabstractApproximate computing (AC) has been considered as a promising paradigm to improve the energy-efficiency of computing hardware for error-tolerant applications, with negligible quality degradation to the output. Dividers frequently limit the performance of a computing system; however, they have not received as much attention as multipliers and adders in AC. In this paper, an energy-efficient and high-performance approximate divider is proposed based on logarithmic conversion and piecewise constant approximation. In this design, the range for the conversion between binary and logarithmic numbers is first expanded from$\mathbf {[{0,1}]}$to$\mathbf {[-0.5,1]}$. A heuristic search algorithm is then devised to find the most accurate constant set to approximate the reciprocal of the divisor, by minimizing a statistical error. The hardware implementation is presented for both floating-point (FP) and integer dividers. With a high configurability, the proposed divider results in a mean relative error distance (MRED) from 2.78% to 0.046%, indicating a high accuracy among state-of-the-art approximate dividers. Compared to the half-precision FP divider, the proposed divider with a MRED of 0.74% can achieve nearly$\mathbf {90\times }$improvement in PDP. Moreover, compared to state-of-the-art approximate dividers, the proposed design is in the Pareto Frontier in terms of power delay product (PDP) and MRED. The three image processing application results demonstrate that the proposed divider can result in the highest peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) even with truncation. Yong Wu 0009, Honglan Jiang, Zining Ma, Pengfei Gou, Jie Han 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2012 | Novel O-GEHL Based Hyperblock Predictor for EDGE ArchitecturesabstractControl flow speculation plays a pushing role to the performance of block-atomic EDGE architectures. Hyperblock predictors, which leverage the style of "Exit + Target" to predict the next hyperblock address, enable high efficient hyperblock-level control flow speculation for EDGE architectures. Recently, a series of binary prediction techniques have been studied and modified to adapt for the exit predictor in hyperblock predictors, including the O-GEHL prediction technique, which was first presented at 1st Championship Branch Prediction Competition. Our paper investigated different mispredict sources in O-GEHL based exit predictor in hyperblock predictors, and proposed two improved strategies: the O-GEHL based exit predictor without chooser and the O-GEHL based exit predictor employing binary O-GEHL prediction. Performance evaluation results showed that: the proposal without chooser outperformed previously published one by 0.7% with the hardware resource ranging from 16KB to 1MB; the proposal employing 8 binary O-GEHL predictor improved the performance by 3% with the largest hardware resource in this paper (1MB); the proposal employing 4 binary O-GEHL predictor for the first 4 exits averagely improved the performance by 2% with the hardware resource ranging from 16KB to 1MB. Pengfei Gou, Mingyan Yu, Zhigang Mao |
NAS | 1 |
| 2010 | M5 based EDGE architecture modelingabstractEDGE (Explicit Data Graph Execution) architectures, a class of architectures distinct from traditional RISC and CISC architectures, have advantages that align well with current technology trends such as power limitations and the need for adaptive exploitation of parallelism. To better understand the architectural and microarchitectural design spaces of EDGE architectures, we have developed a flexible M5-based simulator for EDGE architectures. m5_edge includes a general high-level timing model and ISA support of one specific EDGE ISA. The high-level timing model is not designed to target a specific implementation but common characteristics of EDGE architectures, permitting faster development of a range of microarchitectures. The M5 infrastructure was used because of its high functionality and performance fidelity. The specific EDGE ISA we support is the TRIPS ISA, due to its well-specified ISA and relatively mature compiler. m5_edge can execute binaries generated by the TRIPS toolchain and provides a high-level simulation template for EDGE architectures. Our experimental results show that the difference in execution cycles of m5_edge is within 11% on average, compared to the cycle-accurate simulator provided by the TRIPS group. Thus, m5_edge benefits from both the flexible infrastructure of M5 and acceptable model accuracy, while maintaining both reasonable simulation speed and the ability to quickly explore EDGE-based microarchitectural design spaces. Pengfei Gou, Qingbo Li, Yinghan Jin, Mingyan Yu |
ICCD | 1 |