Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ruiyang Wu 0001

dblp:69/8277 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 since 2021Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 48% Parallel and multicore computing · 23% Electronic design automation · 18%
Software engineering, system software, and programming languages
1 paper
Debugging and program repair · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › debugging support
hardware-assisted deterministic replay
0.322013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
LReplay: a pending period based deterministic replay scheme · ISCA 2010
Processor architecture and microarchitecture
chip multiprocessor
0.212013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Electronic design automation › hardware verification and test
design-for-debug
0.212013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay
0.212013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Memory systems
cache coherence
0.012013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Distributed systems
consistency models
0.012013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Parallel and multicore computing › parallel computing
parallel program debugging
0.012013
Deterministic Replay Using Global Clock · ACM Trans. Archit. Code Optim. 2013
Debugging and program repair › concurrent program debugging
parallel program debugging
0.012010
LReplay: a pending period based deterministic replay scheme · ISCA 2010

Methods — techniques the papers use, named apart from their topics

simulation · 0.2global clock · 0.2direction prediction · 0.2
YearPublicationVenuePosition
2023 SCFM: A Statistical Coarse-to-Fine Method to Select Cross-Microarchitecture Reliable Simulation Points
Chenji Han, Hongze Tan 0001, Xinyu Li 0010, Ruiyang Wu 0001, Fuxin Zhang
APPT5
2015 MRP: mix real cores and pseudo cores for FPGA-based chip-multiprocessor simulation
Xinke Chen, Guangfei Zhang, Huandong Wang, Ruiyang Wu 0001, Longbing Zhang
DATE4
2013 Deterministic Replay Using Global Clock
abstract
Debugging parallel programs is a well-known difficult problem. A promising method to facilitate debugging parallel programs is using hardware support to achieve deterministic replay on a Chip Multi-Processor (CMP). As a Design-For-Debug (DFD) feature, a practical hardware-assisted deterministic replay scheme should have low design and verification costs, as well as a small log size. To achieve these goals, we propose a novel and succinct hardware-assisted deterministic replay scheme named LReplay. The key innovation of LReplay is that instead of recording the logical time orders between instructions or instruction blocks as previous investigations, LReplay is built upon recording the pending period information infused by the global clock. By the recorded pending period information, about 99% execution orders are inferrable, implying that LReplay only needs to record directly the residual 1% noninferrable execution orders in production run. The 1% noninferrable orders can be addressed by a simple yet cost-effective direction prediction technique, which further reduces the log size of LReplay. Benefiting from the preceding innovations, the overall log size of LReplay over SPLASH-2 benchmarks is about 0.17B/K-Inst (byte per k-instruction) for the sequential consistency, and 0.57B/K-Inst for the Godson-3 consistency. Such log sizes are smaller in an order of magnitude than previous deterministic replay schemes incurring no performance loss. Furthermore, LReplay only consumes about 0.5% area of the Godson-3 CMP, since it requires only trivial modifications to existing components of Godson-3. The features of LReplay demonstrate the potential of integrating hardware support for deterministic replay into future industrial processors.
Yunji Chen, Tianshi Chen 0002, Ling Li 0001, Ruiyang Wu 0001, Dao-Fu Liu, Weiwu Hu
ACM Trans. Archit. Code Optim.4
2010 LReplay: a pending period based deterministic replay scheme
abstract
Debugging parallel program is a well-known difficult problem. A promising method to facilitate debugging parallel program is using hardware support to achieve deterministic replay. A hardware-assisted deterministic replay scheme should have a small log size, as well as low design cost, to be feasible for adopting by industrial processors. To achieve the goals, we propose a novel and succinct hardware-assisted deterministic replay scheme named LReplay. The key innovation of LReplay is that instead of recording the logical time orders between instructions or instruction blocks as previous investigations, LReplay is built upon recording the pending period information [6]. According to the experimental results on Godson-3, the overall log size of LReplay is about 0.55B/K-Inst (byte per k-instruction) for sequential consistency, and 0.85B/K-Inst for Godson-3 consistency. The log size is smaller in an order of magnitude than state-of-art deterministic replay schemes incuring no performance loss. Furthermore, LReplay only consumes about $1.3%$ area of Godson-3, since it requires only trivial modifications to the existing components of Godson-3. The above features of LReplay demonstrate the potential of integrating hardware-assisted deterministic replay into future industrial processors.
Yunji Chen, Weiwu Hu, Tianshi Chen 0002, Ruiyang Wu 0001
ISCA4