VLDB 2026 Research / reviewers in the wild / expert
Rishi Khan
dblp:82/68
· DBLP profile ↗
3ranked-venue papers
0as first author
0since 2021 · last 2013
0000-0003-4904-0233ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 61% Distributed systems · 30% Performance modeling and evaluation · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel algorithms › shared-memory parallel algorithms
many-core algorithms |
0.1 | 1 | 2012 | Toward high-throughput algorithms on many-core architectures · ACM Trans. Archit. Code Optim. 2012 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2012 | Toward high-throughput algorithms on many-core architectures · ACM Trans. Archit. Code Optim. 2012 |
Distributed systems
task management |
0.1 | 1 | 2012 | Toward high-throughput algorithms on many-core architectures · ACM Trans. Archit. Code Optim. 2012 |
Performance modeling and evaluation
queueing models |
0.0 | 1 | 2012 | Toward high-throughput algorithms on many-core architectures · ACM Trans. Archit. Code Optim. 2012 |
Methods — techniques the papers use, named apart from their topics
queueing theory · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | A dynamic schema to increase performance in many-core architectures through percolation operationsabstractOptimization of parallel applications under new many-core architectures is challenging even for regular applications. Successful strategies inherited from previous generations of parallel or serial architectures just return incremental gains in performance and further optimization and tuning are required. We argue that conservative static optimizations are not the best fit for modern many-core architectures. The limited advantages of static techniques come from the new scenarios present in many-cores: Plenty of thread units sharing several resources under different coordination mechanisms. Elkin Garcia, Daniel A. Orozco, Rishi Khan, Ioannis E. Venetis, Kelly Livingston, Guang R. Gao |
HiPC | 3 |
| 2012 | Toward high-throughput algorithms on many-core architecturesabstractAdvanced many-core CPU chips already have a few hundreds of processing cores (e.g., 160 cores in an IBM Cyclops-64 chip) and more and more processing cores become available as computer architecture progresses. The underlying runtime systems of such architectures need to efficiently serve hundreds of processors at the same time, requiring all basic data structures within the runtime to maintain unprecedented throughput. In this paper, we analyze the throughput requirements that must be met by algorithms in runtime systems to be able to handle hundreds of simultaneous operations in real time. We reach a surprising conclusion: Many traditional algorithm techniques are poorly suited for highly parallel computing environments because of their low throughput. We reach the conclusion that the intrinsic throughput of a parallel program depends on both its algorithm and the processor architecture where the program runs. We provide theory to quantify the intrinsic throughput of algorithms, and we provide a few examples, where we describe the intrinsic throughput of existing, common algorithms. Then, we go on to explain how to follow a throughput-oriented approach to develop algorithms that have very high intrinsic throughput in many core architectures. We compare our throughput-oriented algorithms with other well known algorithms that provide the same functionality and we show that a throughput-oriented design produces algorithms with equal or faster performance in highly concurrent environments. We provide both theoretical and experimental evidence showing that our algorithms are excellent choices over other state of the art algorithms. The major contributions of this paper are (1) motivating examples that show the importance of throughput in concurrent algorithms; (2) a mathematical framework that uses queueing theory to describe the intrinsic throughput of algorithms; (3) two highly concurrent algorithms with very high intrinsic throughput that are useful for task management in runtime systems; and (4) extensive experimental and theoretical results that show that for highly parallel systems, our proposed algorithms allow greater or at least equal scalability and performance than other well-known similar state-of-the-art algorithms. Daniel A. Orozco, Elkin Garcia, Rishi Khan, Kelly Livingston, Guang R. Gao |
ACM Trans. Archit. Code Optim. | 3 |
| 2010 | Optimized Dense Matrix Multiplication on a Many-Core Architecture
Elkin Garcia, Ioannis E. Venetis, Rishi Khan, Guang R. Gao |
Euro-Par (2) | 3 |