EDBT 2026 Demo / reviewers in the wild / expert
Amik Singh
dblp:72/7751
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 65% Processor architecture and microarchitecture · 14% Performance modeling and evaluation · 10% |
Topics — the 8 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
High-performance computing › performance optimization
many-core processor optimization |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
High-performance computing › numerical linear algebra › linear solver › iterative linear solvers
multigrid method |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
Processor architecture and microarchitecture
SIMD |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUs · PPoPP 2010 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.1 | 1 | 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUs · PPoPP 2010 |
GPUs and heterogeneous computing
GPU computing |
0.0 | 1 | 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUs · PPoPP 2010 |
High-performance computing
sparse linear algebra |
0.0 | 1 | 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUs · PPoPP 2010 |
Methods — techniques the papers use, named apart from their topics
threaded wavefront · 0.1operator fusion · 0.1dynamic threading · 0.1communication aggregation · 0.1performance modeling · 0.1auto-tuning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Optimization of Zuker Algorithm on GPUsabstractPrediction of ribonucleic acid (RNA) secondary structure is one of the most important research areas in bioinformatics. The Zuker algorithm is one of the most popular methods of free energy minimization for RNA secondary structure prediction. We present a novel algorithm to optimize Zuker Algorithm on CUDA GPUs and achieve a speedup of ~10x for certain viruses. Amik Singh, Manoj Misra |
PDCAT | 1 |
| 2012 | Optimization of geometric multigrid for emerging multi- and manycore processorsabstractMultigrid methods are widely used to accelerate the convergence of iterative solvers for linear systems used in a number of different application areas. In this paper, we explore optimization techniques for geometric multigrid on existing and emerging multicore systems including the Opteron-based Cray XE6, Intel® Xeon® E5-2670 and X5550 processor-based Infiniband clusters, as well as the new Intel® Xeon Phi coprocessor (Knights Corner). Our work examines a variety of novel techniques including communication-aggregation, threaded wavefront-based DRAM communication-avoiding, dynamic threading decisions, SIMDization, and fusion of operators. We quantify performance through each phase of the V-cycle for both single-node and distributed-memory experiments and provide detailed analysis for each class of optimization. Results show our optimizations yield significant speedups across a variety of subdomain sizes while simultaneously demonstrating the potential of multi- and manycore processors to dramatically accelerate single-node performance. However, our analysis also indicates that improvements in networks and communication will be essential to reap the potential of manycore processors in large-scale multigrid calculations. Samuel Williams 0001, Dhiraj D. Kalamkar, Amik Singh, Anand M. Deshpande, Brian van Straalen, Mikhail Smelyanskiy, Ann S. Almgren, Pradeep Dubey, John Shalf, Leonid Oliker |
SC | 3 |
| 2011 | Multifrontal Factorization of Sparse SPD Matrices on GPUsabstractSolving large sparse linear systems is often the most computationally intensive component of many scientific computing applications. In the past, sparse multifrontal direct factorization has been shown to scale to thousands of processors on dedicated supercomputers resulting in a substantial reduction in computational time. In recent years, an alternative computing paradigm based on GPUs has gained prominence, primarily due to its affordability, power-efficiency, and the potential to achieve significant speedup relative to desktop performance on regular and structured parallel applications. However, sparse matrix factorization on GPUs has not been explored sufficiently due to the complexity involved in an efficient implementation and concerns of low GPU utilization. In this paper, we present an adaptive hybrid approach for accelerating sparse multifrontal factorization based on a judicious exploitation of the processing power of the host CPU and GPU. We present four different policies for distributing and scheduling the workload between the host CPU and the GPU, and propose a mechanism for a runtime selection of the appropriate policy for each step of sparse Cholesky factorization. This mechanism relies on auto-tuning based on modeling the best policy predictor as a parametric classifier. We estimate the classifier parameters from the available empirical computation time data such that the expected computation time is minimized. This approach is readily adaptable for using the current or an extended set of policies for different CPU-GPU combinations as well as for different combinations of dense kernels for both the CPU and the GPU. Thomas George, Vaibhav Saxena, Amik Singh, Anamitra R. Choudhury |
IPDPS | 4 |
| 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUsabstractWe present a performance model-driven framework for automated performance tuning (autotuning) of sparse matrix-vector multiply (SpMV) on systems accelerated by graphics processing units (GPU). Our study consists of two parts. JeeWhan Choi, Amik Singh, Richard W. Vuduc |
PPoPP | 2 |