Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Vlad Petric

dblp:40/2030 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
1since 2021 · last 2025
0009-0009-6156-9755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 42% Reconfigurable computing and FPGAs · 42% Processor architecture and microarchitecture · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA implementation
0.912025
A Highly-Parallel and Scalable Hardware Accelerator for the NTest Othello Game Engine · IEEE Trans. Parallel Distributed Syst. 2025
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.122005
RENO - A Rename-Based Instruction Optimizer · ISCA 2005
Three extensions to register integration · MICRO 2002
Energy-efficient computing › energy-efficient architecture
energy-efficient microarchitecture
0.112005
Energy-Effectiveness of Pre-Execution and Energy-Aware P-Thread Selection · ISCA 2005
Processor architecture and microarchitecture › speculative execution
pre-execution
0.112005
Energy-Effectiveness of Pre-Execution and Energy-Aware P-Thread Selection · ISCA 2005
Processor architecture and microarchitecture › dynamic optimization
instruction reuse
0.012002
Three extensions to register integration · MICRO 2002
Processor architecture and microarchitecture
speculative execution
0.012002
Three extensions to register integration · MICRO 2002
Processor architecture and microarchitecture › speculative execution
speculative memory bypassing
0.012002
Three extensions to register integration · MICRO 2002
Compilers and program optimization › compiler optimization › redundancy elimination
common subexpression elimination
0.012005
RENO - A Rename-Based Instruction Optimizer · ISCA 2005

Methods — techniques the papers use, named apart from their topics

pattern-based evaluation · 0.9minimax · 0.9alpha-beta pruning · 0.9map-table short-circuiting · 0.1energy simulation · 0.1cycle-level simulation · 0.1critical-path estimation · 0.1simulation · 0.0
YearPublicationVenuePosition
2025 A Highly-Parallel and Scalable Hardware Accelerator for the NTest Othello Game Engine
abstract
Othello is a two-player combinatorial game with 1E+28 legal positions and 1E+58 game tree complexity. We propose a HIghly PArallel, Scalable and configurable hardware accelerator for evaluating the middle and endgame Othello positions. We base HIPAS on NTest - a leading software Othello engine that uses the minimax algorithm with a quality pattern-based evaluation function, alpha-beta pruning, and heuristic mobility sorting. We describe its architecture and Field Programmable Gate Array implementation, measure its performance, and compare it with prior solutions. HIPAS achieves the highest quality evaluation, the highest performance with speed-ups up to several hundreds, and the best energy efficiency. The main novelty is the algorithm implementation as a circular pipeline and a Finite State Machine with pseudo-parallel processing. Although Othello was recently claimed to be weakly solved, the game remains unsolved in a stronger sense. A weak solution only shows how to force a draw. It does not guarantee a win if the opponent makes a mistake. HIPAS can validate the weak solution faster and more efficiently. A multi-threaded NTest software component evaluating the beginning and part of the middle game, combined with one or more instances of HIPAS for handling the remainder can provide a stronger solution.
Stefan Popa, Vlad Petric, Mihai Ivanovici
IEEE Trans. Parallel Distributed Syst.2
2005 Energy-Effectiveness of Pre-Execution and Energy-Aware P-Thread Selection
abstract
Pre-execution removes the microarchitectural latency of "problem" loads from a program's critical path by redundantly executing copies of their computations in parallel with the main program. There have been several proposed pre-execution systems, a quantitative framework (PTHSEL) for analytical pre-execution thread (p-thread) selection, and even a research prototype. To date, however, the energy aspects of pre-execution have not been studied. Cycle-level performance and energy simulations on SPEC2000 integer benchmarks that suffer from L2 misses show that energy-blind pre-execution naturally has a linear latency/energy trade-off, improving performance by 13.8% while increasing energy consumption by 11.9%. To improve this trade-off, we propose two extensions to PTHSEL. First, we replace the flat cycle-for-cycle load cost model with a model based on a critical-path estimation. This extension increases p-thread efficiency in an energy-independent way. Second, we add a parameterized energy model to PTHSEL (forming PTHSEL/sub +E/) that allows it to actively select p-threads that reduce energy rather than (or in combination with) execution latency. Experiments show that PTHSEL/sub +E/ manipulates pre-execution's latency/energy more effectively. Latency targeted selection benefits from the improved load cost model: its performance improvements grow to an average of 16.4% while energy costs drop to 8.7%. ED targeted selection produces p-threads that improve performance by only 12.9%, but ED by 8.8%. Targeting p-thread selection for energy reduction, results in "energy-free" pre-execution, with average speedup of 5.4%, and a small decrease in total energy consumption (0.7%).
Vlad Petric, Amir Roth
ISCA1
2005 RENO - A Rename-Based Instruction Optimizer
abstract
RENO is a modified MIPS R10000 register renamer that uses map-table "short-circuiting" to implement dynamic versions of several well-known static optimizations: move elimination, common subexpression elimination, register allocation, and constant folding. Because it implements these optimizations dynamically, RENO can apply optimizations in certain situations where static compilers cannot. Cycle-level simulation shows that RENO dynamically eliminates (i.e. optimizes away) 22% of the dynamic instructions in both SPECint2000 and MediaBench. RENO/sub CF/ is responsible for 12% and 17% of the eliminations, respectively. Because dataflow dependences are collapsed around eliminated instructions, performance improves by 8% and 13%, respectively. Alternatively, because eliminated instructions do not consume issue queue entries, physical registers, or issue, bypass, register file, and execution bandwidth, RENO can be used to absorb the performance impact of a significantly scaled-down execution core.
Vlad Petric, Tingting Sha, Amir Roth
ISCA1
2002 Three extensions to register integration
abstract
Register integration (or just integration) is a register renaming discipline that implements instruction reuse via physical register sharing. Initially developed to perform squash reuse, the integration mechanism can exploit more reuse scenarios. Here, we describe three extensions to the original design that expand its applicability and boost its performance impact. First, we extend squash reuse to general reuse. Whereas squash reuse maintains the concept of an instruction instance "owning" its output register, we allow multiple instructions to simultaneously share a single register. Next, we replace the PC-indexing scheme with an opcode-based indexing scheme that exposes more integration opportunities. Finally, we introduce an extension called reverse integration in which we speculatively create integration entries for the inverses of operations for instance, when renaming an add, we create an entry for the inverse subtract. Reverse integration allows us to reuse operations that the program itself has not executed yet. We use reverse integration to implement speculative memory bypassing for stack-pointer based loads (register fills and restores). Our evaluation shows that these extensions increase the integration rate - the number of retired instructions that integrate older results and bypass the execution engine -to an average of 15% on the SPEC2000 integer benchmarks. On a 4-way superscalar processor with an aggressive memory system, this translates into an average IPC improvement of 7%. The fact that integrating instructions completely bypass the execution engine raises the possibility of using integration as a low-complexity substitute for execution bandwidth and issue buffering. Our experiments show that such a trade-off is possible, enabling a range of IPC/complexity designs.
Vlad Petric, Anne Bracy, Amir Roth
MICRO1