Shivam Kundan

dblp:273/5836 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-1733-6358ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% Memory systems · 50%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › task scheduling
contention-aware scheduling
0.612022
A Pressure-Aware Policy for Contention Minimization on Multicore Systems · ACM Trans. Archit. Code Optim. 2022
Operating systems › resource management › process management › CPU scheduling
thread scheduling
0.212022
A Pressure-Aware Policy for Contention Minimization on Multicore Systems · ACM Trans. Archit. Code Optim. 2022

Methods — techniques the papers use, named apart from their topics

performance monitoring counters · 1.1cache monitoring technology · 1.1
YearPublicationVenuePosition
2022 A Pressure-Aware Policy for Contention Minimization on Multicore Systems
abstract
Modern Chip Multiprocessors (CMPs) are integrating an increasing amount of cores to address the continually growing demand for high-application performance. The cores of a CMP share several components of the memory hierarchy, such as Last-Level Cache (LLC) and main memory. This allows for considerable gains in multithreaded applications while also helping to maintain architectural simplicity. However, sharing resources can also result in performance bottleneck due to contention among concurrently executing applications. In this work, we formulate a fine-grained application characterization methodology that leverages Performance Monitoring Counters (PMCs) and Cache Monitoring Technology (CMT) in Intel processors. We utilize this characterization methodology to develop two contention-aware scheduling policies, one static and one dynamic , that co-schedule applications based on their resource-interference profiles. Our approach focuses on minimizing contention on both the main-memory bandwidth and the LLC by monitoring the pressure that each application inflicts on these resources. We achieve performance benefits for diverse workloads, outperforming Linux and three state-of-the-art contention-aware schedulers in terms of system throughput and fairness for both single and multithreaded workloads. Compared with Linux, our policy achieves up to 16% greater throughput for single-threaded and up to 40% greater throughput for multithreaded applications. Additionally, the policies increase fairness by up to 65% for single-threaded and up to 130% for multithreaded ones.
Shivam Kundan, Theodoros Marinakis, Iraklis Anagnostopoulos, Dimitrios Kagaris
ACM Trans. Archit. Code Optim.1
2021 Priority-Aware Scheduling under Shared-Resource Contention on Chip Multicore Processors
abstract
In this paper, we present a priority-aware scheduling methodology for concurrent application execution on chip multi-core processors. Our methodology improves the performance of up to 4 high-priority applications while also preventing resource starvation for low-priority ones by means of progress-aware scheduling. We compare our results with Linux's completely fair scheduler and two state-of-the-art progress aware schedulers. Experimental results on an Intel Xeon Gold 6130 CPU demonstrate an average increase in high-priority application performance of 36.4% over Linux while also maintaining high throughput for low-priority applications. In addition, we show that our methodology achieves high-priority application performance comparable to a state-of-the-art hardware cache partitioning method (Intel's POCAT). Our method achieves average performance within 14.5% to 0.2% of POCAT without the need for hardware support.
Shivam Kundan, Iraklis Anagnostopoulos
ISCAS1
2020 A Machine Learning Approach for Improving Power Efficiency on Clustered Multi-Processor System
abstract
Modern embedded systems have adopted the clustered Chip Multi-Processor (CMP) paradigm in conjunction with dynamic frequency scaling techniques for improving application performance and power consumption. However, applications suffer from performance saturation due to resource contention. Thus, after a certain point, any frequency increase results only in power overheads without any performance gains. In this work, we present a run-time manager that focuses on power efficiency improvement for clustered CMPs. Specifically, it monitors the activity of concurrently executing applications and utilizes neural networks to select an appropriate frequency that keeps performance high, while reducing power consumption. Experimental results on the Odroid-XU3 board show that the proposed methodology improves power efficiency (MIPS/Watt) by 23% compared to Linux's performance governor with a negligible performance drop of only 3%.
Shivam Kundan, Iraklis Anagnostopoulos
ISCAS1