Viacheslav V. Fedorov

dblp:155/4384 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-2916-7694ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 56% Integrated circuit design · 42% Energy-efficient computing · 2%
Computer networks
1 paper
Routing and switching · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
content-addressable memory
0.822021
Hardware Acceleration of Hash Operations in Modern Microprocessors · IEEE Trans. Computers 2021
FTCAM: An Area-Efficient Flash-Based Ternary CAM Design · IEEE Trans. Computers 2016
Integrated circuit design
digital circuit design
0.512021
Hardware Acceleration of Hash Operations in Modern Microprocessors · IEEE Trans. Computers 2021
Integrated circuit design › digital arithmetic circuits
special function unit
0.512021
Hardware Acceleration of Hash Operations in Modern Microprocessors · IEEE Trans. Computers 2021
Memory systems › content-addressable memory
TCAM
0.212016
FTCAM: An Area-Efficient Flash-Based Ternary CAM Design · IEEE Trans. Computers 2016
Memory systems › cache management
cache replacement
0.212013
ARI: Adaptive LLC-memory traffic management · ACM Trans. Archit. Code Optim. 2013
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.212013
ARI: Adaptive LLC-memory traffic management · ACM Trans. Archit. Code Optim. 2013
Memory systems
non-volatile memory
0.212013
ARI: Adaptive LLC-memory traffic management · ACM Trans. Archit. Code Optim. 2013
Memory systems
write-back reduction
0.212013
ARI: Adaptive LLC-memory traffic management · ACM Trans. Archit. Code Optim. 2013
Routing and switching
IP lookup
0.112016
FTCAM: An Area-Efficient Flash-Based Ternary CAM Design · IEEE Trans. Computers 2016
Routing and switching
routing
0.112016
FTCAM: An Area-Efficient Flash-Based Ternary CAM Design · IEEE Trans. Computers 2016
Energy-efficient computing
memory system energy
0.012013
ARI: Adaptive LLC-memory traffic management · ACM Trans. Archit. Code Optim. 2013

Methods — techniques the papers use, named apart from their topics

microarchitecture simulation · 0.5circuit simulation · 0.5SPICE simulation · 0.5adaptive replacement and insertion · 0.2
YearPublicationVenuePosition
2021 Hardware Acceleration of Hash Operations in Modern Microprocessors
abstract
Modern microprocessors contain several special function units (SFUs) such as specialized arithmetic units, cryptographic processors, etc. In recent times, applications such as cloud computing, web-based search engines, and network applications are widely used, and place new demands on the microprocessor. Hashing is a key algorithm that is extensively used in such applications. Hashing can reduce the complexity of search and lookup from O(N) to O(N/n), where n bins are used. Hashing is typically performed in software. Thus, implementing a hardware-based hash unit on a modern microprocessor would potentially increase performance significantly. In this article, we propose a novel hardware hash unit (HU) design for use in modern microprocessors, at the microarchitecture level and at the circuit level. First, we present the design of the HU at the microarchitecture level. We simulate the HU to compare its performance with a software-based hash implementation. We demonstrate a significant speedup (up to 15×) for the HU. Furthermore, the performance scales elegantly with increasing database size and application diversity, without increasing the hardware cost. Second, we present the circuit design of the HU for use in modern microprocessors, using a 45nm technology. Our proposed hardware hash unit is based on the use of a content-addressable memory (CAM) to implement each bin of the hash table. We simulate the HU circuit and compare it with a traditional CAM design. We demonstrate an average power reduction of 5.48× using the HU over the traditional CAM. Also, we show that the HU can operate at a maximum frequency of 1.39 GHz (after accounting for process, voltage and temperature (PVT) variations and accounting for wiring parasitics). Furthermore, we present the delay, power and area trade-offs of the HU design with varying hash table sizes.
Abbas A. Fairouz, Monther Abusultan, Viacheslav V. Fedorov, Sunil P. Khatri
IEEE Trans. Computers3
2016 FTCAM: An Area-Efficient Flash-Based Ternary CAM Design
abstract
This paper presents a Ternary Content-addressable Memory (TCAM) design which is based on the use of floating-gate (flash) transistors. TCAMs are extensively used in high speed IP networking, and are commonly found in routers in the internet core. Traditional TCAM ICs are built using CMOS devices, and a single TCAM cell utilizes 17 transistors. In contrast, our TCAM cell utilizes only two flash transistors, thereby significantly reducing circuit area. We cover the chip-level architecture of the TCAM IC briefly, focusing mainly on the TCAM block which does fast parallel IP routing table lookup. Our flash-based TCAM (FTCAM) block is simulated in SPICE, and we show that it has a significantly lowered area compared to a CMOS based TCAM block, with a speed that can meet current ($\sim$400 Gb/s) data rates that are found in the internet core.
Viacheslav V. Fedorov, Monther Abusultan, Sunil P. Khatri
IEEE Trans. Computers1
2014 An area-efficient Ternary CAM design using floating gate transistors
abstract
This paper presents a Ternary Content-addressable Memory (TCAM) design which is based on the use of floating-gate (flash) transistors. TCAMs are extensively used in high speed IP networking, and are commonly found in routers in the internet core. Traditional TCAM ICs are built using CMOS devices, and a single TCAM cell utilizes 17 transistors. In contrast, our TCAM cell utilizes only 2 flash transistors, thereby significantly reducing circuit area. We cover the chip-level architecture of the TCAM IC briefly, focusing mainly on the TCAM block which does fast parallel IP routing table lookup. Our flash based TCAM block is simulated in SPICE, and we show that it has a significantly lowered area compared to a CMOS based TCAM block, with a speed that can meet current (~400 Gb/s) data rates that are found in the internet core.
Viacheslav V. Fedorov, Monther Abusultan, Sunil P. Khatri
ICCD1
2013 ARI: Adaptive LLC-memory traffic management
abstract
Decreasing the traffic from the CPU LLC to main memory is a very important issue in modern systems. Recent work focuses on cache misses, overlooking the impact of writebacks on the total memory traffic, energy consumption, IPC, and so forth. Policies that foster a balanced approach, between reducing write traffic to memory and improving miss rates, can increase overall performance and improve energy efficiency and memory system lifetime for NVM memory technology, such as phase-change memory (PCM). We propose Adaptive Replacement and Insertion (ARI), an adaptive approach to last-level CPU cache management, optimizing the two parameters (miss rate and writeback rate) simultaneously. Our specific focus is to reduce writebacks as much as possible while maintaining or improving the miss rate relative to conventional LRU replacement policy. ARI reduces LLC writebacks by 33%, on average, while also decreasing misses by 4.7%, on average. In a typical system, this boosts IPC by 4.9%, on average, while decreasing energy consumption by 8.9%. These results are achieved with minimal hardware overheads.
Viacheslav V. Fedorov, Sheng Qiu, A. L. Narasimha Reddy, Paul Gratz
ACM Trans. Archit. Code Optim.1