EDBT 2026 Demo / reviewers in the wild / expert
Vahid Janfaza
dblp:137/9153
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-8656-4848ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Cost-Effective Dueling Framework for Set-Associative Cache IndexingabstractPathological program behavior may cause a non-uniform access distribution in set associative caches, leading to an increase in conflict misses.To address this challenge, prior works profile the program patterns and propose different index functions to avoid these conflict misses [11,18].However, as we analyze the prior work on set-associative cache indexing, we identify two major issues.First, there is no single index function that is guaranteed to perform well for every application.Second, advanced indexing schemes typically have sophisticated implementation and prohibitively long computation latency.In this paper, we propose Duelhash, a dynamic N-way indexing framework for set associative caches, which provides an effective dueling mechanism for multiple index functions at runtime with a simple and efficient hardware implementation.At runtime, the performance of the index functions are evaluated periodically, and the best performer is applied to the cache.To evaluate the performance of Duelhash, we conduct a case study on a 16-way set-associative LLC using a diverse set of benchmarks, including SPEC 2006, SPEC 2017, PARSEC 3.0, CVP and GAP.Our empirical results show that without prefetching, Duelhash provides an IPC speed up of 2.8% (with the highest being 23%) over the conventional powerof-two modulo (Default) index, compared to a 1.6% speed up of a commercialized indexing scheme (Xorhash).When pattern-based prefetchers are turned on in the L1 data and L2 caches, Duelhash can provide up to 5.8% single-core speedup over Default.Duelhash also provides a 6.2% MPKI Kevin Weston, Vahid Janfaza, Avery Johnson, Abdullah Muzahid |
ICS | 2 |
| 2024 | Customizing Cache Indexing Through Entropy EstimationabstractModern computers heavily rely on caches as one of the means to achieve higher performance. As a result, cache management has been the topic of extensive research. Compared to cache replacement and prefetching, cache indexing has re-ceived far less interest over the years. Being in the critical path, a good cache index function must exhibit a high performance while having a minimal computational delay. Previous indexing schemes fall short of these requirements, having either moderate performance or a prohibitively expensive delay. We propose ENTROPyINDEX, an entropy-based cache indexing scheme that can deliver superior performance while maintaining a minimal computational cost. ENTROPyINDEX is based on the idea of constructing the index function dynamically at runtime using the address bits with the highest entropy (randomness) to maximize the balance of the cache access distribution. The entropy of the address bits is measured by determining which bits change between two subsequent cache misses. ENTROPyINDEX periodically compares the entropy of different bits and selects the ones that change the most. This dynamic selection scheme allows ENTROPyINDEX to adapt to different types of applications. Our experimental results show that ENTROPyINDEX outper-forms previous indexing schemes both with and without hardware prefetching. For SPEC 2006, SPEC 2017, PARSEC 3.0 and GAP benchmarks without prefetching, ENTROPyINDEX delivers a geometric mean IPC improvement of 3.39% (with the highest being 52.2%), compared to a 1.74% improvement of the state-of-the-art index function (PRIME) and a 1.76% improvement of a commercialized indexing scheme (XORHASH) over the baseline power-of-two modulo scheme. With prefetching, ENTROPyINDEX is the only indexing scheme with a substantial performance gain of 1.42% (with the highest being 30.1 %), compared to a 0.41 % improvement of Prime and a 0.49% improvement of Xorhash over the same baseline. For non-uniform applications and no-prefetching, ENTROPyINDEX gives an IPC speed up of 5.58%, compared to a 2.26% speed up of Prime and a 2.23% speed up of Xorhash. For non-uniform applications with prefetching, the IPC speed up of ENTROPyINDEX is 2.08%, compared to a 0.35% speed up of Prime and a 0.53% speed up of Xorhash. For CVP workloads without prefetching, ENTROPyINDEX delivers a speed up of 3.04% over the baseline compared to a 1.52% of Prime and a 2.04% of Xorhash. For CVP workloads with prefetching, ENTROPyINDEX improves the IPC by 1.60%, compared to 0.63% of Prime and 1.07% of Xorhash. Kevin Weston, Avery Johnson, Vahid Janfaza, Farabi Mahmud, Abdullah Muzahid |
MICRO | 3 |
| 2023 | MERCURY: Accelerating DNN Training By Exploiting Input SimilarityabstractDeep Neural Networks (DNN) are computationally intensive to train. It consists of a large number of multidimensional dot products between many weights and input vectors. However, there can be significant similarities among input vectors. If one input vector is similar to another, its computations with the weights are similar to those of the other and, therefore, can be skipped by reusing the already-computed results. We propose a novel scheme, called MERCURY, to exploit input similarity during DNN training in a hardware accelerator. MERCURY uses Random Projection with Quantization (RPQ) to convert an input vector to a bit sequence, called Signature. A cache (MCACHE) stores signatures of recent input vectors along with the computed results. If the Signature of a new input vector matches that of an already existing vector in the MCACHE, the two vectors are found to have similarities. Therefore, the already-computed result is reused for the new vector. To the best of our knowledge, MERCURY is the first work that exploits input similarity using RPQ for accelerating DNN training in hardware. The paper presents a detailed design, workflow, and implementation of the MERCURY. Our experimental evaluation with twelve different deep learning models shows that MERCURY saves a significant number of computations and speeds up the model training by an average of 1.97× with an accuracy similar to the baseline system. Vahid Janfaza, Kevin Weston, Moein Razavi, Shantanu Mandal, Farabi Mahmud, Alex Hilty, Abdullah Muzahid |
HPCA | 1 |
| 2023 | ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionabstractNeural network training is inherently sequential where the layers finish the forward propagation in succession, followed by the calculation and back-propagation of gradients (based on a loss function) starting from the last layer. The sequential computations significantly slow down neural network training, especially the deeper ones. Prediction has been successfully used in many areas of computer architecture to speed up sequential processing. Therefore, we propose ADA-GP, which uses gradient prediction adaptively to speed up deep neural network (DNN) training while maintaining accuracy. ADA-GP works by incorporating a small neural network to predict gradients for different layers of a DNN model. ADA-GP uses a novel tensor reorganization method to make it feasible to predict a large number of gradients. ADA-GP alternates between DNN training using backpropagated gradients and DNN training using predicted gradients. ADA-GP adaptively adjusts when and for how long gradient prediction is used to strike a balance between accuracy and performance. Last but not least, we provide a detailed hardware extension in a typical DNN accelerator to realize the speed up potential from gradient prediction. Our extensive experiments with fifteen DNN models show that ADA-GP can achieve an average speed up of 1.47 × with similar or even higher accuracy than the baseline models. Moreover, it consumes, on average, 34% less energy due to reduced off-chip memory accesses compared to the baseline accelerator. Vahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah Muzahid |
MICRO | 1 |