VLDB 2026 Research / reviewers in the wild / expert
Amitava Majumdar 0001
dblp:20/5122 · also Amit Majumdar 0001
· DBLP profile ↗
11ranked-venue papers
6as first author
0since 2021 · last 2015
0000-0002-0860-6686ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Electronic design automation · 60% Processor architecture and microarchitecture · 24% Performance modeling and evaluation · 12% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
hardware verification and test |
0.0 | 2 | 1996 | On Evaluating and Optimizing Weights for Weighted Random Pattern Testing · IEEE Trans. Computers 1996 Fault Coverage and Test Length Estimation for Random Pattern Testing · IEEE Trans. Computers 1995 |
Electronic design automation › hardware verification and test
test length estimation |
0.0 | 2 | 1996 | On Evaluating and Optimizing Weights for Weighted Random Pattern Testing · IEEE Trans. Computers 1996 Fault Coverage and Test Length Estimation for Random Pattern Testing · IEEE Trans. Computers 1995 |
Processor architecture and microarchitecture
memory latency tolerance |
0.0 | 1 | 1998 | Multi-processor Performance on the Tera MTA · SC 1998 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 1998 | Multi-processor Performance on the Tera MTA · SC 1998 |
Performance modeling and evaluation
parallel performance evaluation |
0.0 | 1 | 1998 | Multi-processor Performance on the Tera MTA · SC 1998 |
Electronic design automation › hardware verification and test › random testing
weighted random pattern testing |
0.0 | 1 | 1996 | On Evaluating and Optimizing Weights for Weighted Random Pattern Testing · IEEE Trans. Computers 1996 |
Electronic design automation › hardware verification and test
fault coverage |
0.0 | 1 | 1995 | Fault Coverage and Test Length Estimation for Random Pattern Testing · IEEE Trans. Computers 1995 |
Electronic design automation › hardware verification and test
random testing |
0.0 | 1 | 1995 | Fault Coverage and Test Length Estimation for Random Pattern Testing · IEEE Trans. Computers 1995 |
Methods — techniques the papers use, named apart from their topics
amdahl's law analysis · 0.0NAS benchmarks · 0.0hill-climbing · 0.0convex optimization · 0.0approximation bounds · 0.0sequential sampling · 0.0probability generating function · 0.0moment analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Early experiences in developing and managing the neuroscience gatewayabstractThe last few decades have seen the emergence of computational neuroscience as a mature field where researchers are interested in modeling complex and large neuronal systems and require access to high performance computing machines and associated cyber infrastructure to manage computational workflow and data. The neuronal simulation tools, used in this research field, are also implemented for parallel computers and suitable for high performance computing machines. But using these tools on complex high performance computing machines remains a challenge because of issues with acquiring computer time on these machines located at national supercomputer centers, dealing with complex user interface of these machines, dealing with data management and retrieval. The Neuroscience Gateway is being developed to alleviate and/or hide these barriers to entry for computational neuroscientists. It hides or eliminates, from the point of view of the users, all the administrative and technical barriers and makes parallel neuronal simulation tools easily available and accessible on complex high performance computing machines. It handles the running of jobs and data management and retrieval. This paper shares the early experiences in bringing up this gateway and describes the software architecture it is based on, how it is implemented, and how users can use this for computational neuroscience research using high performance computing at the back end. We also look at parallel scaling of some publicly available neuronal models and analyze the recent usage data of the neuroscience gateway. Subhashini Sivagnanam, Amitava Majumdar 0001, Kenneth Yoshimoto, Vadim Astakhov, Anita E. Bandrowski, Maryann E. Martone, Nicholas T. Carnevale |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | XSEDE13 Special Issue Conference PublicationsabstractThis special issue is based on presentations at the XSEDE13 conference held in San Diego, July 2013.The conference is an annual gathering of the extended community of those interested in advancing research cyberinfrastructure (CI) and digital resources.This includes researchers using XSEDE and like resources, XSEDE staff members, Campus Champions, and especially students.Seven hundred twenty people attended the conference, from 50 states, 1 territory, and 14 countries.As a summertime conference with terrific student programs from high school to post-graduate levels, we were fortunate to attract 200 students.From the Call for Participation, the purpose of the conference is to 'showcases the discoveries, innovations, challenges and achievements of those who utilize and support XSEDE resources and services, as well as other digital resources and services throughout the world'.The conference included tutorials, birds-of-a-feather sessions, a poster session, visualization showcase, and job fair.Each conference has a unique theme.In 2013, the focus was the impact of science gateways and the relevance of computation and high-end analysis in the biosciences.The core technical program consisted of four tracks-science and engineering; technology; software and software environments; and training, education, and outreach.Selected papers from the core technical program were invited to expand content and submit to this special issue. SCIENCE AND ENGINEERING TRACKXSEDE13's science track received 47 submissions, 27 of which were accepted for the conference.Eight of the best were asked to develop extended papers for inclusion in this special journal issue.Papers in the science track covered various research topics and all of them involve utilizing XSEDE resources in innovative ways.Research topics span wide areas such as biological sequence analysis, modeling of energy usage in buildings, analysis of error correcting codes (ECC) for simulations on GPUs, the impact of Campus Bridging for researchers transitioning to XSEDE resources, algorithms for computational finance analyses, methods enabling improved search capabilities over large digitized document archives, and the implementation of algorithms on Intel Xeon Phi coprocessors.The paper on functional annotation of newly sequenced genomes [1] describes an optimized workflow to enable large-scale protein annotation.It utilizes a special classification algorithm and high performance computing (HPC), and the demonstrated results show capabilities which scientists will be able to utilize to annotate big genome data.The Building Energy Modeling (BEM) approach is combined with machine learning methods in another paper [2] to enable efficient modeling of US buildings.The parametric space makes the use of supercomputers a necessity, and the results are utilized to train machine learning algorithms.Another paper [3] looks at ECC available on many of the modern graphics processing units (GPUs) used in HPC machines.It looks at the penalty involved in utilizing ECC and compares molecular dynamics simulation results regarding ECC events triggered during such simulations.It discusses if error checking is necessary for such simulations given the penalty associated with it.Transitioning from campus resources to XSEDE resources is not an easy task for researchers.Making Campus Bridging Work for Researchers [4] proposes utilization of Campus Bridging experts to make this transition easier while requiring minimal investment from the organizing body. Nancy Wilkins-Diehr, Amitava Majumdar 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Quantifying performance benefits of overlap using MPI-2 in a seismic modeling applicationabstractAWM-Olsen is a widely used ground motion simulation code based on a parallel finite difference solution of the 3-D velocity-stress wave equation. This application runs on tens of thousands of cores and consumes several million CPU hours on the TeraGrid Clusters every year. A significant portion of its run-time (37% in a 4,096 process run), is spent in MPI communication routines. Hence, it demands an optimized communication design coupled with a low-latency, high-bandwidth network and an efficient communication subsystem for good performance. In this paper, we analyze the performance bottlenecks of the application with regard to the time spent in MPI communication calls. We find that much of this time can be overlapped with computation using MPI non-blocking calls. We use both two-sided and MPI-2 one-sided communication semantics to re-design the communication in AWM-Olsen. We find that with our new design, using MPI-2 one-sided communication semantics, the entire application can be sped up by 12% at 4K processes and by 10% at 8K processes on a state-of-the-art InfiniBand cluster, Ranger at the Texas Advanced Computing Center (TACC). Sreeram Potluri, Ping Lai, Karen A. Tomko, Sayantan Sur, Yifeng Cui, Mahidhar Tatineni, Karl W. Schulz, William L. Barth, Amitava Majumdar 0001, Dhabaleswar K. Panda 0001 |
ICS | 9 |
| 2000 | arallel Performance Study of Monte Carlo Photon Transport Code on Shared-, Distributed-, and Distributed-Shared-Memory ArchitecturesabstractWe have parallelized a Monte Carlo photon transport algorithm. Three different parallel versions of the algorithm were developed. The first version is for the Tera Multi-Threaded Architecture (MTA) and uses Tera specific directives. The second version, which uses MPI library calls, has been implemented on both the CRAY T3E and the 8-way SMP IBM SP with Power3 processors. The third version is a hybrid MPI-OpenMP implementation and is used on the SMP IBM SP. This version uses MPI to communicate between nodes and OpenMP to perform shared memory operations among processors within a node. We explain the three different parallelization approaches and present parallel performance results of these three parallel implementations on three different machines. We observe near perfect speedup for the three versions on the three architectures. The results on the SMP IBM SP suggest that the hybrid MPI-OpenMP programming is suitable for SMP type machines. Amitava Majumdar 0001 |
IPDPS | 1 |
| 1998 | Multi-processor Performance on the Tera MTAabstractThe Tera MTA is a revolutionary commercial computer based on a multithreaded processor architecture. In contrast to many other parallel architectures, the Tera MTA can effectively use high amounts of parallelism on a single processor. By running multiple threads on a single processor, it can tolerate memory latency and to keep the processor saturated. If the computation is sufficiently large, it can benefit from running on multiple processors. A primary architectural goal of the MTA is that it provide scalable performance over multiple processors. This paper is a preliminary investigation of the first multi-processor Tera MTA. In a previous paper [1] we reported that on the kernel NAS 2 benchmarks [2], a single-processor MTA system running at the architected clock speed would be similar in performance to a single processor of the Cray T90. We found that the compilers of both machines were able to find the necessary threads or vector operations, after making standard changes to the random number generator. In this paper we update the single-processor results in two ways: we use only actual clock speeds, and we report improvements given by further tuning of the MTA codes. We then investigate the performance of the best single-processor codes when run on a two-processor MTA, making no further tuning effort. The parallel efficiency of the codes range from 77% to 99%. An analysis shows that the "serial bottlenecks" -- unparallelized code sections and the cost of allocating and freeing the parallel hardware resources -- account for less than a percent of the runtimes. Thus, Amdahl's Law needn't take effect on the NAS benchmarks until there are hundreds of processors running thousands of threads. Instead, the major source of inefficiency appears to be an imperfect network connecting the processors to the memory. Ideally, the network can support one memory reference per instruction. The current hardware has defects that reduce the throughput to about 85% of this rate. Except for the EP benchmark, the tuned codes issue memory references at nearly the peak rate of one per instruction. Consequently, the network can support the memory references issued by one, but not two, processors. As a result, the parallel efficiency of EP is near- perfect, but the others are reduced accordingly. Another reason for imperfect speedup pertains to the compiler. While the definition of a thread in a single processor or multi-processor mode is essentially the same, there is a different implementation and an associated overhead with running on multiple processors. We characterize the overhead of running "frays" (a collection of threads running on a single processor) and "crews" (a collection of frays, one per processor.) Allan Snavely, Larry Carter, Jay Boisseau, Amitava Majumdar 0001, Kang Su Gatlin, Nick Mitchell, John Feo, Brian D. Koblenz |
SC | 4 |
| 1996 | On Evaluating and Optimizing Weights for Weighted Random Pattern TestingabstractTwo problems in weighted random pattern testing are considered: 1) evaluating a set of input weights in terms of the amount of time required to generate a set of test patterns and 2) determining the optimal weights for a given test set. An exact expression for expected test length is derived as a function of input weights. Upper and lower bounds for expected test length are presented. Percentage error of approximation is expressed in terms of the bounds. Based on these results, algorithms are given for approximating expected test length. These algorithms allow the user to tradeoff accuracy and computational complexity. Experiments with some test sets are presented to illustrate the accuracy of the approximation technique. Expected test length is shown to be a convex function of input weights. A simple hill-climbing algorithm is defined to find optimal weights for a given set of test patterns. When hardware constraints, limiting the number of weights to be realized for each input bit, are also specified, a simple modification of the algorithm suffices in yielding optimal weights in the constrained space. Experiments with several circuits yield of the order of 96% to 99% reduction in expected test length over that achieved by current techniques. Amitava Majumdar 0001 |
IEEE Trans. Computers | 1 |
| 1995 | Fault Coverage and Test Length Estimation for Random Pattern TestingabstractFault coverage and test length estimation in circuits under random test is the subject of this paper. Testing by a sequence of random input patterns is viewed as sequential sampling of faults from a given fault universe. Based on this model, the probability mass function (pmf) of fault coverage and expressions for all its moments are derived. This provides a means for computing estimates of fault coverage as well as determining the accuracy of the estimates. Test length, viewed as waiting time on fault coverage, is analyzed next. We derive expressions for its pmf and its probability generating function (pgf). This allows computation of all the higher order moments. In particular, expressions for mean and variance of test length for any specified fault coverage are derived. This is a considerable enhancement of the state of the art in techniques for predicting test length as a function of fault coverage. It is shown that any moment of test length requires knowledge of all the moments of fault coverage, and hence, its pmf. For this reason, expressions for approximating its expected value and variance, for user specified error bounds, are also given. A methodology based on these results is outlined. Experiments carried out on several circuits demonstrate that this technique is capable of providing excellent predictions of test length. Furthermore it is shown, as with fault coverage prediction, that estimates of variances can be used to bound average test length quite effectively.> Amitava Majumdar 0001, Sarma B. K. Vrudhula |
IEEE Trans. Computers | 1 |
| 1994 | WRAPTure: A Tool for Evaluation and Optimization of Weights for Weighted Random Pattern TestingabstractTwo problems in weighted random pattern testing are addressed: 1) finding the expected test length to sample all patterns in a given test set for a given weight set; and 2) finding the optimal weight set for minimizing expected test length for a given test set. Exact analytical expressions for expected test length for sampling all patterns from a test set are given. Expressions for approximating this quantity are also derived. Based on these results an algorithm is formulated to obtain the optimal weight assignment. Experiments carried out on test sets for several circuits (including ISCAS '85 benchmarks) yield of the order of 95% to 100% reduction in expected test length over that achieved by current techniques.> Amitava Majumdar 0001 |
ICCD | 1 |
| 1994 | Techniques for estimating test length under random test
Amitava Majumdar 0001, Sarma B. K. Vrudhula |
J. Electron. Test. | 1 |
| 1994 | On Lookahead in the List Update Problem
Rahul Simha, Amitava Majumdar 0001 |
Inf. Process. Lett. | 2 |
| 1993 | Analysis of signal probability in logic circuits using stochastic modelsabstractAnalyzes the behavior of signal probabilities in logic circuits chosen from a statistically characterized population. The statistical parameters of the population are obtained from certain aggregate structural and logical characteristics of the circuit such as fanins, fanouts, and proportions of different types of gates. A circuit is first transformed into one consisting of only nand gates, inverters, and buffers. This transformation leads to a new classification of circuits, referred to as nor-type, or-type, nand-type and and-type, the particular type being determined by computing two parameters from the circuit specification. A functional relation between gate signal probabilities, primary input signal probabilities, and aggregate structural properties of a circuit is established. This allows the study of basic characteristics of signal probability and its limiting behavior when the number of levels increases. It is shown that the limiting behavior of signal probability depends on the fixed points of a function which is determined by the two parameters estimated from the circuit and the distribution of gate fanins. A recurrence relation also allows one to define a methodology for estimating the distribution of signal probabilities in different levels. The complexity of this technique is shown to be proportional to the number of levels in the circuit. Results of extensive experiments with ISCAS '85 benchmarks as well as other circuits are given.> Amitava Majumdar 0001, Sarma B. K. Vrudhula |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |