Ilya Ganusov

dblp:80/4703 · also Ilya K. Ganusov · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 66% Processor architecture and microarchitecture · 20% Performance modeling and evaluation · 9%
Software engineering, system software, and programming languages
2 papers
Program analysis · 54% Runtime systems and virtual machines · 46%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA architecture
0.412020
Architectural Enhancements in Intel® Agilex™ FPGAs · FPGA 2020
Reconfigurable computing and FPGAs
FPGA routing architecture
0.412020
Architectural Enhancements in Intel® Agilex™ FPGAs · FPGA 2020
Processor architecture and microarchitecture
value prediction
0.122006
The VPC Trace-Compression Algorithms · IEEE Trans. Computers 2005
Future execution: A prefetching mechanism that uses multiple cores to speed up single threads · ACM Trans. Archit. Code Optim. 2006
Processor architecture and microarchitecture
multicore design
0.112006
Future execution: A prefetching mechanism that uses multiple cores to speed up single threads · ACM Trans. Archit. Code Optim. 2006
Memory systems › cache
prefetching
0.112006
Future execution: A prefetching mechanism that uses multiple cores to speed up single threads · ACM Trans. Archit. Code Optim. 2006
Program analysis
trace compression
0.112005
The VPC Trace-Compression Algorithms · IEEE Trans. Computers 2005
Performance modeling and evaluation
trace compression
0.112005
The VPC Trace-Compression Algorithms · IEEE Trans. Computers 2005
Runtime systems and virtual machines
binary translation
0.012004
Automatic Synthesis of High-Speed Processor Simulators · MICRO 2004
Performance modeling and evaluation
simulation
0.012004
Automatic Synthesis of High-Speed Processor Simulators · MICRO 2004
Performance modeling and evaluation
trace analysis
0.012005
The VPC Trace-Compression Algorithms · IEEE Trans. Computers 2005

Methods — techniques the papers use, named apart from their topics

time borrowing · 0.4clock skew · 0.4value prediction · 0.1compression algorithm design · 0.1profile-guided optimization · 0.1mixed-mode execution · 0.1runahead execution · 0.1hardware stream prefetching · 0.1future execution · 0.1
YearPublicationVenuePosition
2021 Clock Skew Scheduling: Avoiding the Runtime Cost of Mixed-Integer Linear Programming
abstract
Clock Skew Scheduling has become a common practice in state-of-the-art FPGAs with the introduction of delay chains on the clock path in the hardware of both Xilinx and Intel®FPGAs, as well as clock skew scheduling algorithms in the CAD tools. Ideally, globally optimal solutions are sought to find the best solution across the entire design. However, using Mixed-Integer Linear Programming (MILP) to find such optimal solutions has a large and sometimes unrealistic runtime cost, especially for bigger designs. Besides, a high runtime does not necessarily correlate to improved performance. We present, in this paper, techniques to reduce the runtime of mixed-integer linear programming approaches to clock skew scheduling, and provide alternatives that can achieve optimal or near-optimal performance gain with a fraction of the MILP runtime cost.
Grace Zgheib, Yu Shen Lu, Ilya Ganusov
FPL3
2020 Architectural Enhancements in Intel® Agilex™ FPGAs
abstract
This paper describes architectural enhancements in Intel® Agilex™ FPGAs and SoCs. Agilex devices are built on Intel's 10nm process and feature next-generation programmable fabric, tightly coupled with a quad-core ARM processor subsystem, a secure device manager, IO and memory interfaces, and multiple companion transceiver tile choices. The Agilex fabric features multiple logic block enhancements that significantly improve propagation delays and integrate more effectively with the second-generation HyperFlexAgilex™ pipelined routing architecture. Routing connections are re-designed to be point-to-point, dropping intermediate connections featured in prior FPGA generations and replacing them with a wider variety of shorter wire types. Fine-grain programmable clock skew and time-borrowing were introduced throughout the fabric to augment the slack-balancing capabilities of HyperFlex registers. DSP capabilities are also extended to natively support new INT9/BFLOAT16/FP16 formats. Together, along with process and circuit enhancements, these changes support more than 40% performance improvement over the Stratix® 10 family of FPGAs.
Jeffrey Chromczak, Mark Wheeler, Charles Chiasson, Dana How, Martin Langhammer, Tim Vanderhoek, Grace Zgheib, Ilya Ganusov
FPGA8
2020 Agilex™ Generation of Intel® FPGAs
abstract
This article consists only of a collection of slides from the author's conference presentation.
Ilya Ganusov, Mahesh A. Iyer, Alon Meisler
Hot Chips Symposium1
2016 Time-borrowing platform in the Xilinx UltraScale+ family of FPGAs and MPSoCs
abstract
This paper presents enhancements to the Xilinx UltraScale+ clocking architecture to support fine-grain time-borrowing. Time borrowing improves performance by redistributing timing slack between fast and slow paths. The Ultra-Scale+ architecture introduces programmable hardware delays and pulse generators embedded in the clocking tree to support time-borrowing based both on clock skew scheduling and pulsed latches. This programmable hardware allows borrowing from a few picoseconds to multiple nanoseconds between sequential pipeline stages without any changes to RTL, placement or routing. Vivado algorithms automatically determine when to skew flip-flop clock or convert them to pulsed latches to achieve the highest possible performance. Using the default Vivado flow, this programmable time-borrowing platform delivers 5.5% Fmaxincrease on average over a suite of 89 industrial designs. It is especially effective on high-speed applications, delivering up to 13.7% Fmaxincrease on individual designs. We also demonstrate that using non-default features, such as delays cascades or increasing hold margin, can increase average performance gains to 7.4% and 8.5%, respectively. This platform incurs minimum area (less than 0.1% of total chip area) while staying robust in the presence of tight hold constraints and increasing process variation.
Ilya Ganusov, Benjamin Devlin
FPL1
2016 Automated extra pipeline analysis of applications mapped to Xilinx UltraScale+ FPGAs
abstract
This paper describes the methodology and algorithms behind extra pipeline analysis tools released in the Xilinx Vivado Design Suite version 2015.3. Extra pipelining is one of the most effective ways to improve performance of FPGA applications. Manual pipelining, however, often requires significant efforts from FPGA designers who need to explore various changes in the RTL and re-run the flow iteratively. The automatic pipelining approach described in this paper, in contrast, allows FPGA users to explore latency vs. performance trade-offs of their designs before investing time and effort into modifying RTL. We describe algorithms behind these tools which use simple cut heuristics to maximize performance improvement while minimizing additional latency and register overhead. To demonstrate the effectiveness of the proposed approach, we analyse a set of 93 commercial FPGA applications and IP blocks mapped to Xilinx UltraScale+ and UltraScale generations of FPGAs. The results show that extra pipelining can provide from 18% to 29% potential Fmax improvement on average. It also shows that the distribution of improvements is bimodal, with almost half of benchmark suite designs showing no improvement due to the presence of large loops. Finally, we demonstrate that highly-pipelined designs map well to UltraScale+ and UltraScale FPGA architectures. Our approach demonstrates 19% and 20% Fmax improvement potential for the UltraScale+ and UltraScale architectures respectively, with the majority of applications reaching their loop limit through pipelining.
Ilya Ganusov, Henri Fraisse, Aaron Ng, Rafael Trapani Possignolo, Sabya Das
FPL1
2015 UltraScale+ MPSoC and FPGA families
Vamsi Boppana, Sagheer Ahmad, Ilya Ganusov, Vinod Kathail, Vidya Rajagopalan, Ralph Wittig
Hot Chips Symposium3
2006 Efficient emulation of hardware prefetchers via event-driven helper threading
abstract
The advance of multi-core architectures provides significant benefits for parallel and throughput-oriented computing, but the performance of individual computation threads does not improve and may even suffer a penalty because of the increased contention for shared resources. This paper explores the idea of using available general-purpose cores in a CMP as helper engines for individual threads running on the active cores. We propose a lightweight architectural framework for efficient event-driven software emulation of complex hardware accelerators and describe how this framework can be applied to implement a variety of prefetching techniques. We demonstrate the viability and effectiveness of our framework on a wide range of applications from the SPEC CPU2000 and Olden benchmark suites. On average, our mechanism provides performance benefits within 5% of pure hardware implementations. Furthermore, we demonstrate that running event-driven prefetching threads on top of a baseline with a hardware stride prefetcher yields significant speedups for many programs. Finally, we show that our approach provides competitive performance improvements over other hardware approaches for multi-core execution while executing fewer instructions and requiring considerably less hardware support.
Ilya Ganusov, Martin Burtscher
PACT1
2006 Future execution: A prefetching mechanism that uses multiple cores to speed up single threads
abstract
This paper describes future execution (FE), a simple hardware-only technique to accelerate individual program threads running on multicore microprocessors. Our approach uses available idle cores to prefetch important data for the threads executing on the active cores. FE is based on the observation that many cache misses are caused by loads that execute repeatedly and whose address-generating program slices do not change (much) between consecutive executions. To exploit this property, FE dynamically creates a prefetching thread for each active core by simply sending a copy of all committed, register-writing instructions to an otherwise idle core. The key innovation is that on the way to the second core, a value predictor replaces each predictable instruction in the prefetching thread with a load immediate instruction, where the immediate is the predicted result that the instruction is likely to produce during its n th next dynamic execution. Executing this modified instruction stream (i.e., the prefetching thread) on another core allows to compute the future results of the instructions that are not directly predictable, issue prefetches into the shared memory hierarchy, and thus reduce the primary threads' memory access time. We demonstrate the viability and effectiveness of future execution by performing cycle-accurate simulations of a two-way CMP running the single-threaded SPECcpu2000 benchmark suite. Our mechanism improves program performance by 12%, on average, over a baseline that already includes an optimized hardware stream prefetcher. We further show that FE is complementary to runahead execution and that the combination of these two techniques raises the average speedup to 20% above the performance of the baseline processor with the aggressive stream prefetcher.
Ilya Ganusov, Martin Burtscher
ACM Trans. Archit. Code Optim.1
2005 The VPC Trace-Compression Algorithms
abstract
Execution traces, such as are used to study and analyze program behavior, are often so large that they need to be stored in compressed form. This paper describes the design and implementation of four value prediction-based compression (VPC) algorithms for traces that record the PC as well as other information about executed instructions. VPC1 directly compresses traces using value predictors, VPC2 adds a second compression stage, and VPC3 utilizes value predictors to convert traces into streams that can be compressed better and more quickly than the original traces. VPC4 introduces further algorithmic enhancements and is automatically synthesized. Of the 55 SPECcpu2000 traces we evaluate, VPC4 compresses 36 better, decompresses 26 faster, and compresses 53 faster than BZIP2, MACHE, PDATS II, SBC, and SEQUITUR. It delivers the highest geometric-mean compression rate, decompression speed, and compression speed because of the predictors' simplicity and their ability to exploit local value locality. Most other compression algorithms can only exploit global value locality.
Martin Burtscher, Ilya Ganusov, Sandra J. Jackson, Jian Ke, Paruj Ratanaworabhan, Nana B. Sam
IEEE Trans. Computers2
2004 Automatic Synthesis of High-Speed Processor Simulators
abstract
Microprocessor simulators are very popular in research and teaching environments. For example, functional simulators are often used to perform architectural studies, to fast-forward over uninteresting code, to generate program traces, and to warm up tables before switching to a more detailed but slower simulator. Unfortunately, most portable functional simulators are on the order of 100 times slower than native execution. This paper describes a set of novel techniques and optimizations to synthesize portable functional simulators that are only 6.6 times slower on average (16 times in the worst case) than native execution and 19 times faster than SimpleScalar's sim-fast on the SPECcpu2000 programs. When simulating a memory hierarchy, the synthesized code is 2.6 times faster than the equivalent ATOM code. Our fully automated synthesis approach works without access to source/assembly code or debug information. It generates C code, integrates optional user-provided code, performs unwanted-code removal, preserves basic blocks, generates low-overhead profiles, employs a simple heuristic to determine potential jump targets, only compiles important instructions, and utilizes mixed-mode execution, i.e., it interleaves compiled and interpreted simulation to maximize performance.
Martin Burtscher, Ilya Ganusov
MICRO2