Akbar Sharifi

dblp:74/288 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 7 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 52% Performance modeling and evaluation · 12% Processor architecture and microarchitecture · 11%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.222012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011
Processor architecture and microarchitecture
multicore design
0.222011
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011
Process variation-aware routing in NoC based multicores · DAC 2011
Memory systems › cache management
cache capacity management
0.112012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Memory systems
memory access latency
0.112012
Addressing End-to-End Memory Access Latency in NoC-Based Multicores · MICRO 2012
Memory systems
memory latency reduction
0.112012
Addressing End-to-End Memory Access Latency in NoC-Based Multicores · MICRO 2012
Memory systems › cache management
shared cache management
0.112012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Electronic design automation › physical design › routing › message routing
network-on-chip routing
0.112011
Process variation-aware routing in NoC based multicores · DAC 2011
Cloud and datacenter computing › resource management
shared resource management
0.112011
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011
Performance modeling and evaluation › workload characterization
multiprogrammed workloads
0.012012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Parallel and multicore computing › thread-level parallelism
multithreaded workloads
0.012012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Hardware reliability and fault tolerance
process variation
0.012011
Process variation-aware routing in NoC based multicores · DAC 2011
Parallel and multicore computing
resource coordination
0.012011
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011

Methods — techniques the papers use, named apart from their topics

response prioritization · 0.1request prioritization · 0.1priority-based thread scheduling · 0.1source routing algorithm · 0.1resource management · 0.1qos scheduling · 0.1
YearPublicationVenuePosition
2017 DEMM: A Dynamic Energy-Saving Mechanism for Multicore Memories
abstract
Since main memory system contributes to a large and increasing fraction of server/datacenter energy consumption, there have been several efforts to reduce its power and energy consumption. DVFS schemes have been used to reduce the memory power, but they come with a performance penalty. In this work, we propose DEMM, an OS-based, high performance DVFS mechanism that reduces memory power by dynamically scaling individual memory channel frequencies/voltages. Our strategy also involves clustering the running applications based on their sensitivities to memory latency, and assigning memory channels to the application clusters. We introduce a new metric called Discrete Misses per Kilo Cycle (DMPKC) to capture the performance sensitivities of the applications to memory frequency modulation. DEMM allows us to save power in the memory system with negligible impact on performance. We demonstrate around 25% savings in the memory system energy and 10% savings in the total system energy, with only a 4% loss in workload performance.
Akbar Sharifi, Wei Ding 0008, Diana R. Guttman, Hui Zhao 0013, Xulong Tang, Mahmut T. Kandemir, Chita R. Das
MASCOTS1
2012 PEPON: performance-aware hierarchical power budgeting for NoC based multicores
abstract
Targeting NoC based multicores, we propose a two-level power budget distribution mechanism, called PEPON, where the first level distributes the overall power budget of the multicore system among various types of on-chip resources like the cores, caches, and NoC, and the second level determines the allocation of power to individual instances of each type of resource. Both these distributions are oriented towards maximizing workload performance without exceeding the specified power budget. Extensive experimental evaluations of the proposed power distribution scheme using a full system simulation and detailed power models emphasize the importance of power budget partitioning at both levels. Specifically, our results show that the proposed scheme can provide up to 29% performance improvement as compared to no power budgeting, and performs 13% better than a competing scheme, under the same chip-wide power cap.
Akbar Sharifi, Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das
PACT1
2012 Courteous cache sharing: being nice to others in capacity management
abstract
This paper proposes a cache management scheme for multiprogrammed, multithreaded applications, with the objective of obtaining maximum performance for both individual applications and the multithreaded workload mix. In this scheme, each individual application's performance is improved by increasing the priority of its slowest thread, while the overall system performance is improved by ensuring that each individual application's performance benefit does not come at the cost of a significant degradation to other application's threads that are sharing the same cache. Averaged over six workloads, our shared cache management scheme improves the performance of the combination of applications by 18%. These improvements across applications in each mix are also fair, as indicated by average fair speedup improvements of 10% across the threads of each application (averaged over all the workloads).
Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir, Mary Jane Irwin
DAC1
2012 Addressing End-to-End Memory Access Latency in NoC-Based Multicores
abstract
To achieve high performance in emerging multicores, it is crucial to reduce the number of memory accesses that suffer from very high latencies. However, this should be done with care as improving latency of an access can worsen the latency of another as a result of resource sharing. Therefore, the goal should be to balance latencies of memory accesses issued by an application in an execution phase, while ensuring a low average latency value. Targeting Network-on-Chip (NoC) based multicores, we propose two network prioritization schemes that can cooperatively improve performance by reducing end-to-end memory access latencies. Our first scheme prioritizes memory response messages such that, in a given period of time, messages of an application that experience higher latencies than the average message latency for that application are expedited and a more uniform memory latency pattern is achieved. Our second scheme prioritizes the request messages that are destined for idle memory banks over others, with the goal of improving bank utilization and preventing long queues from being built in front of the memory banks. These two network prioritization-based optimizations together lead to uniform memory access latencies with a low average value. Our experiments with a 4×8 mesh network-based multicore show that, when applied together, our schemes can achieve 15%, 10% and 13% performance improvement on memory intensive, memory non-intensive, and mixed multiprogrammed workloads, respectively.
Akbar Sharifi, Emre Kultursay, Mahmut T. Kandemir, Chita R. Das
MICRO1
2011 Process variation-aware routing in NoC based multicores
abstract
We propose a variation-aware source routing algorithm for a heterogenous NoC where each router has a different operating latency, as a result of process variations. Our proposed scheme computes the best path for each communication, based on the inherent speed of the routers (dictated by process variations) and the current traffic pattern. Our results indicate that employing our proposed routing scheme reduces average packet latencies (our performance metric), in our applications, by up to 28% as compared to the deterministic and adaptive routing algorithms.
Akbar Sharifi, Mahmut T. Kandemir
DAC1
2011 Feedback control based cache reliability enhancement for emerging multicores
abstract
Focusing on data reliability, we propose a control theory centric approach designed to improve transient error resilience in shared caches of emerging multicores while satisfying performance goals. The proposed scheme takes, as input, two quality of service (QoS) specifications: performance QoS and reliability QoS. The first of these indicates the minimum workload-wide cache (L2) hit rate value acceptable, whereas the second one captures the reliability bound on an application basis, with the help of a metric called the Reads-with-Replica (RwR). We present an extensive experimental evaluation of the proposed scheme on various workloads formed using the applications from the SPEC2006 benchmark suite. The proposed scheme is able to satisfy, in most of the tested cases, both performance and reliability QoS targets, by successfully modulating the total size of the data replication area and partitioning of this area among the co-runner applications. The collected results also show that our scheme achieves consistent improvements under different values of the major simulation parameters.
Hui Zhao 0013, Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir
ICCAD2
2011 Automatic Feedback Control of Shared Hybrid Caches in 3D Chip Multiprocessors
abstract
3D integration enables building caches from different types of technologies such as SRAM, Magnetic RAM (MRAM), DRAM, and Phase-change RAM (PRAM). Hybrid cache architectures (HCAs) have been proposed to take advantage of the benefits offered by these types of technologies. Employing this novel cache architecture to build shared caches in chip multiprocessors (CMPs) can lead to significant performance and power consumption improvements. In this paper, we focus on a 3D CMP design in which the shared last level L2 cache is composed of an MRAM layer and an SRAM layer stacked upon the processing cores and present a control theory centric approach designed to partition this shared hybrid L2 cache space dynamically among concurrently running applications in order to satisfy the application-level performance QoS targets. At each time interval, the two layers of the hybrid L2 cache are partitioned, based on the cache demands made by the controllers of the applications, to satisfy the specified performance targets. We evaluate our feedback control based scheme using various workloads. Our experimental evaluation shows that the proposed scheme is able to satisfy the specified performance QoS in most of the tested cases, by partitioning the hybrid cache space of the 3D CMP among co-runner applications.
Akbar Sharifi, Mahmut T. Kandemir
PDP1
2011 METE: meeting end-to-end QoS in multicores through system-wide resource management
abstract
Management of shared resources in emerging multicores for achieving predictable performance has received considerable attention in recent times. In general, almost all these approaches attempt to guarantee a certain level of performance QoS (weighted IPC, harmonic speedup, etc) by managing a single shared resource or at most a couple of interacting resources. A fundamental shortcoming of these approaches is the lack of coordination between these shared resources to satisfy a system level QoS. This is undesirable because providing end-to-end QoS in future multicores is essential for supporting wide-spread adoption of these architectures in virtualized servers and cloud computing systems. An initial step towards such an end-to-end QoS support in multicores is to ensure that at least the major computational and memory resources on-chip are managed efficiently in a coordinated fashion.
Akbar Sharifi, Shekhar Srikantaiah, Asit K. Mishra, Mahmut T. Kandemir, Chita R. Das
SIGMETRICS1
2010 Feedback control for providing QoS in NoC based multicores
abstract
In this paper, we employ formal feedback control theory to achieve desired communication throughput across a network-on-chip (NoC) based multicore. When the output of the system needs to follow a certain reference input over time, our controller regulates the system to obtain the desired effect on the output. In this work, targeting a multicore that executes multiple applications simultaneously, we demonstrate how to design and employ a PID (Proportional Integral Derivative) controller to obtain the desired throughput for communications by tuning the weights of the virtual channels of the routers in the NoC. We also propose a global controller architecture that implements policies to handle situations in which the network cannot provide the overlapping communications with sufficient resources or the throughputs of the communications can be enhanced (beyond their specified values) due to the availability of excess resources. Finally, we discuss how our novel control architecture works under different scenarios by presenting experimental results obtained using four embedded applications. These results show how the global controller adjusts the virtual channels weights to achieve the desired throughputs of different communications across the NoC, and as a result, the system output successfully tracks the specified input.
Akbar Sharifi, Hui Zhao 0013, Mahmut T. Kandemir
DATE1