Khubaib

dblp:64/8277 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 52% Memory systems · 44% Energy-efficient computing · 3%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
chip multiprocessor
0.322012
MorphCore: An Energy-Efficient Microarchitecture for High Performance ILP and High Throughput TLP · MICRO 2012
Data marshaling for multi-core architectures · ISCA 2010
Memory systems
cache
0.212016
Accelerating Dependent Cache Misses with an Enhanced Memory Controller · ISCA 2016
Memory systems › cache
cache miss
0.212016
Accelerating Dependent Cache Misses with an Enhanced Memory Controller · ISCA 2016
Processor architecture and microarchitecture
branch prediction
0.212013
Store-Load-Branch (SLB) predictor: A compiler assisted branch prediction for data dependent branches · HPCA 2013
Processor architecture and microarchitecture › out-of-order execution
out-of-order core
0.112012
MorphCore: An Energy-Efficient Microarchitecture for High Performance ILP and High Throughput TLP · MICRO 2012
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.112012
MorphCore: An Energy-Efficient Microarchitecture for High Performance ILP and High Throughput TLP · MICRO 2012
Memory systems
cache management
0.112010
Data marshaling for multi-core architectures · ISCA 2010
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication
0.112010
Data marshaling for multi-core architectures · ISCA 2010
Memory systems
memory controller
0.112016
Accelerating Dependent Cache Misses with an Enhanced Memory Controller · ISCA 2016
Parallel and multicore computing
parallel programming models
0.012010
Data marshaling for multi-core architectures · ISCA 2010

Methods — techniques the papers use, named apart from their topics

dynamic branch prediction · 0.3prefetching · 0.2global history buffer · 0.2core fusion · 0.1profiling · 0.1data marshaling · 0.1
YearPublicationVenuePosition
2016 Accelerating Dependent Cache Misses with an Enhanced Memory Controller
abstract
On-chip contention increases memory access latency for multi-core processors. We identify that this additional latency has a substantial effect on performance for an important class of latency-critical memory operations: those that result in a cache miss and are dependent on data from a prior cache miss. We observe that the number of instructions between the first cache miss and its dependent cache miss is usually small. To minimize dependent cache miss latency, we propose adding just enough functionality to dynamically identify these instructions at the core and migrate them to the memory controller for execution as soon as source data arrives from DRAM. This migration allows memory requests issued by our new Enhanced Memory Controller (EMC) to experience a 20% lower latency than if issued by the core. On a set of memory intensive quad-core workloads, the EMC results in a 13% improvement in system performance and a 5% reduction in energy consumption over a system with a Global History Buffer prefetcher, the highest performing prefetcher in our evaluation.
Milad Hashemi, Khubaib, Eiman Ebrahimi, Onur Mutlu, Yale N. Patt
ISCA2
2013 Store-Load-Branch (SLB) predictor: A compiler assisted branch prediction for data dependent branches
abstract
Data-dependent branches constitute single biggest source of remaining branch mispredictions. Typically, data-dependent branches are associated with program data structures, and follow store-load-branch execution sequence. A set of memory locations is written at an earlier point in a program. Later, these locations are read, and used for evaluating branch condition. Branch outcome depends on data values stored in data structure, which, typically do not have repeatable pattern. Therefore, in addition to history-based dynamic predictor, we need a different kind of predictor for handling such branches. This paper presents Store-Load-Branch (SLB) predictor; a compiler-assisted dynamic branch prediction scheme for data-dependent direct and indirect branches. For every data-dependent branch, compiler identifies store instructions that modify the data structure associated with the branch. Marked store instructions are dynamically tracked, and stored values are used for computing branch flags ahead of time. Branch flags are buffered, and later used for making predictions. On average, compared to standalone TAGE predictor, combined TAGE+SLB predictor reduces branch MPKI by 21% and 51% for SPECINT and EEMBC benchmark suites respectively.
Muhammad Umar Farooq 0003, Khubaib, Lizy Kurian John
HPCA2
2012 MorphCore: An Energy-Efficient Microarchitecture for High Performance ILP and High Throughput TLP
abstract
Several researchers have recognized in recent years that today's workloads require a micro architecture that can handle single-threaded code at high performance, and multi-threaded code at high throughput, while consuming no more energy than is necessary. This paper proposes Morph Core, a unique approach to satisfying these competing requirements, by starting with a traditional high performance out-of-order core and making minimal changes that can transform it into a highly-threaded in-order SMT core when necessary. The result is a micro architecture that outperforms an aggressive 4-way SMT out-of-order core, "medium" out-of-order cores, small in-order cores, and Core Fusion. Compared to a 2-way SMT out-of-order core, Morph Core increases performance by 10% and reduces energy-delay-squared product by 22%.
Khubaib, M. Aater Suleman, Milad Hashemi, Chris Wilkerson, Yale N. Patt
MICRO1
2012 Energy Savings via Dead Sub-Block Prediction
abstract
Cache memories have traditionally been designed to exploit spatial locality by fetching entire cache lines from memory upon a miss. However, recent studies have shown that often the number of sub-blocks within a line that are actually used is low. Furthermore, those sub-blocks that are used are accessed only a few times before becoming dead (i.e., never accessed again). This results in considerable energy waste since 1) data not needed by the processor is brought into the cache, and 2) data is kept alive in the cache longer than necessary. We propose the Dead Sub-Block Predictor (DSBP) to predict which sub-blocks of a cache line will be actually used and how many times it will be used in order to bring into the cache only those sub-blocks that are necessary, and power them off after they are touched the predicted number of times. We also use DSBP to identify dead lines (i.e., all sub-blocks off) and augment the existing replacement policy by prioritizing dead lines for eviction. Our results show a 24% energy reduction for the whole cache hierarchy when averaged over the SPEC2000, SPEC2006 and NAS-NPB benchmarks.
Marco A. Z. Alves, Khubaib, Eiman Ebrahimi, Veynu Narasiman, Carlos Villavieja, Philippe Olivier Alexandre Navaux, Yale N. Patt
SBAC-PAD2
2010 Feedback-directed pipeline parallelism
abstract
Extracting high performance from Chip Multiprocessors requires that the application be parallelized. A common software technique to parallelize loops is pipeline parallelism in which the programmer/compiler splits each loop iteration into stages and each stage runs on a certain number of cores. It is important to choose the number of cores for each stage carefully because the core-to-stage allocation determines performance and power consumption. Finding the best core-to-stage allocation for an application is challenging because the number of possible allocations is large, and the best allocation depends on the input set and machine configuration.
M. Aater Suleman, Moinuddin K. Qureshi, Khubaib, Yale N. Patt
PACT3
2010 Data marshaling for multi-core architectures
abstract
Previous research has shown that Staged Execution (SE), i.e., dividing a program into segments and executing each segment at the core that has the data and/or functionality to best run that segment, can improve performance and save power. However, SE's benefit is limited because most segments access inter-segment data, i.e., data generated by the previous segment. When consecutive segments run on different cores, accesses to inter-segment data incur cache misses, thereby reducing performance. This paper proposes Data Marshaling (DM), a new technique to eliminate cache misses to inter-segment data. DM uses profiling to identify instructions that generate inter-segment data, and adds only 96 bytes/core of storage overhead. We show that DM significantly improves the performance of two promising Staged Execution models, Accelerated Critical Sections and producer-consumer pipeline parallelism, on both homogeneous and heterogeneous multi-core systems. In both models, DM can achieve almost all of the potential of ideally eliminating cache misses to inter-segment data. DM's performance benefit increases with the number of cores.
M. Aater Suleman, Onur Mutlu, José A. Joao, Khubaib, Yale N. Patt
ISCA4