EDBT 2026 Demo / reviewers in the wild / expert
Ruiyang Wu 0001
dblp:69/8277
· DBLP profile ↗
4ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 48% Parallel and multicore computing · 23% Electronic design automation · 18% | |
| Software engineering, system software, and programming languages
1 paper |
Debugging and program repair · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › debugging support
hardware-assisted deterministic replay |
0.3 | 2 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 LReplay: a pending period based deterministic replay scheme · ISCA 2010 |
Processor architecture and microarchitecture
chip multiprocessor |
0.2 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Electronic design automation › hardware verification and test
design-for-debug |
0.2 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay |
0.2 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Memory systems
cache coherence |
0.0 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Distributed systems
consistency models |
0.0 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Parallel and multicore computing › parallel computing
parallel program debugging |
0.0 | 1 | 2013 | Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013 |
Debugging and program repair › concurrent program debugging
parallel program debugging |
0.0 | 1 | 2010 | LReplay: a pending period based deterministic replay scheme · ISCA 2010 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2global clock · 0.2direction prediction · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SCFM: A Statistical Coarse-to-Fine Method to Select Cross-Microarchitecture Reliable Simulation Points
Chenji Han, Hongze Tan 0001, Xinyu Li 0010, Ruiyang Wu 0001, Fuxin Zhang |
APPT | 5 |
| 2015 | MRP: mix real cores and pseudo cores for FPGA-based chip-multiprocessor simulation
Xinke Chen, Guangfei Zhang, Huandong Wang, Ruiyang Wu 0001, Longbing Zhang |
DATE | 4 |
| 2013 | Deterministic Replay Using Global ClockabstractDebugging parallel programs is a well-known difficult problem. A promising method to facilitate debugging parallel programs is using hardware support to achieve deterministic replay on a Chip Multi-Processor (CMP). As a Design-For-Debug (DFD) feature, a practical hardware-assisted deterministic replay scheme should have low design and verification costs, as well as a small log size. To achieve these goals, we propose a novel and succinct hardware-assisted deterministic replay scheme named LReplay. The key innovation of LReplay is that instead of recording the logical time orders between instructions or instruction blocks as previous investigations, LReplay is built upon recording the pending period information infused by the global clock. By the recorded pending period information, about 99% execution orders are inferrable, implying that LReplay only needs to record directly the residual 1% noninferrable execution orders in production run. The 1% noninferrable orders can be addressed by a simple yet cost-effective direction prediction technique, which further reduces the log size of LReplay. Benefiting from the preceding innovations, the overall log size of LReplay over SPLASH-2 benchmarks is about 0.17B/K-Inst (byte per k-instruction) for the sequential consistency, and 0.57B/K-Inst for the Godson-3 consistency. Such log sizes are smaller in an order of magnitude than previous deterministic replay schemes incurring no performance loss. Furthermore, LReplay only consumes about 0.5% area of the Godson-3 CMP, since it requires only trivial modifications to existing components of Godson-3. The features of LReplay demonstrate the potential of integrating hardware support for deterministic replay into future industrial processors. Yunji Chen, Tianshi Chen 0002, Ling Li 0001, Ruiyang Wu 0001, Dao-Fu Liu, Weiwu Hu |
ACM Trans. Archit. Code Optim. | 4 |
| 2010 | LReplay: a pending period based deterministic replay schemeabstractDebugging parallel program is a well-known difficult problem. A promising method to facilitate debugging parallel program is using hardware support to achieve deterministic replay. A hardware-assisted deterministic replay scheme should have a small log size, as well as low design cost, to be feasible for adopting by industrial processors. To achieve the goals, we propose a novel and succinct hardware-assisted deterministic replay scheme named LReplay. The key innovation of LReplay is that instead of recording the logical time orders between instructions or instruction blocks as previous investigations, LReplay is built upon recording the pending period information [6]. According to the experimental results on Godson-3, the overall log size of LReplay is about 0.55B/K-Inst (byte per k-instruction) for sequential consistency, and 0.85B/K-Inst for Godson-3 consistency. The log size is smaller in an order of magnitude than state-of-art deterministic replay schemes incuring no performance loss. Furthermore, LReplay only consumes about $1.3%$ area of Godson-3, since it requires only trivial modifications to the existing components of Godson-3. The above features of LReplay demonstrate the potential of integrating hardware-assisted deterministic replay into future industrial processors. Yunji Chen, Weiwu Hu, Tianshi Chen 0002, Ruiyang Wu 0001 |
ISCA | 4 |