EDBT 2026 Demo / reviewers in the wild / expert
Manish Gupta 0002
dblp:g/ManishGupta2 · also Manish K. Gupta 0002
· DBLP profile ↗
32ranked-venue papers
6as first author
1since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-authorSoftware engineering, systems software and programming languages · 10 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
High-performance computing · 48% Parallel and multicore computing · 12% Memory systems · 10% | |
| Software engineering, system software, and programming languages
10 papers |
Compilers and program optimization · 34% Runtime systems and virtual machines · 26% Program analysis · 17% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › supercomputing
bluegene/l |
0.1 | 3 | 2006 | Unlocking the Performance of the BlueGene/L Supercomputer · SC 2004 An overview of the BlueGene/L Supercomputer · SC 2002 High performance file I/O for the Blue Gene/L supercomputer · HPCA 2006 |
Runtime systems and virtual machines
garbage collection |
0.1 | 2 | 2002 | Exploiting prolific types for memory management and optimizations · POPL 2002 Creating and preserving locality of java applications at allocation and garbage collection times · OOPSLA 2002 |
Operating systems › resource management
memory management |
0.1 | 2 | 2002 | Exploiting prolific types for memory management and optimizations · POPL 2002 Creating and preserving locality of java applications at allocation and garbage collection times · OOPSLA 2002 |
Program analysis › static analysis › pointer analysis
escape analysis |
0.1 | 2 | 2003 | Stack allocation and synchronization optimizations for Java using escape analysis · ACM Trans. Program. Lang. Syst. 2003 Escape Analysis for Java · OOPSLA 1999 |
Concurrent programming
lock elimination |
0.1 | 2 | 2003 | Stack allocation and synchronization optimizations for Java using escape analysis · ACM Trans. Program. Lang. Syst. 2003 Escape Analysis for Java · OOPSLA 1999 |
High-performance computing
finite element method |
0.1 | 1 | 2006 | Gordon Bell finalists I - Large scale drop impact analysis of mobile phone using ADVC on Blue Gene/L · SC 2006 |
High-performance computing
parallel i/o |
0.1 | 1 | 2006 | High performance file I/O for the Blue Gene/L supercomputer · HPCA 2006 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2006 | Gordon Bell finalists I - Large scale drop impact analysis of mobile phone using ADVC on Blue Gene/L · SC 2006 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2006 | Gordon Bell finalists I - Large scale drop impact analysis of mobile phone using ADVC on Blue Gene/L · SC 2006 |
High-performance computing
structural analysis |
0.1 | 1 | 2006 | Gordon Bell finalists I - Large scale drop impact analysis of mobile phone using ADVC on Blue Gene/L · SC 2006 |
Compilers and program optimization › parallel program optimization
synchronization optimization |
0.1 | 2 | 2003 | Stack allocation and synchronization optimizations for Java using escape analysis · ACM Trans. Program. Lang. Syst. 2003 Static Analysis to Reduce Synchronization Costs in Data-Parallel Programs · POPL 1996 |
High-performance computing
supercomputing |
0.1 | 2 | 2006 | An overview of the BlueGene/L Supercomputer · SC 2002 High performance file I/O for the Blue Gene/L supercomputer · HPCA 2006 |
Integrated circuit design › digital circuit design › arithmetic circuit design
floating-point unit |
0.0 | 1 | 2004 | Unlocking the Performance of the BlueGene/L Supercomputer · SC 2004 |
High-performance computing
supercomputer architecture |
0.0 | 1 | 2004 | Unlocking the Performance of the BlueGene/L Supercomputer · SC 2004 |
Compilers and program optimization › parallel program optimization
communication optimization |
0.0 | 3 | 1996 | A Unified Framework for Optimizing Communication in Data-Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 Global Communication Analysis and Optimization · PLDI 1996 An HPF Compiler for the IBM SP2 · SC 1995 |
Program analysis
static analysis |
0.0 | 1 | 2003 | Stack allocation and synchronization optimizations for Java using escape analysis · ACM Trans. Program. Lang. Syst. 2003 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2003 | Critical event prediction for proactive management in large-scale computer clusters · KDD 2003 |
Distributed systems › fault tolerance
proactive fault prediction |
0.0 | 1 | 2003 | Critical event prediction for proactive management in large-scale computer clusters · KDD 2003 |
Runtime systems and virtual machines
object representation |
0.0 | 1 | 2002 | Exploiting prolific types for memory management and optimizations · POPL 2002 |
Memory systems › data locality
cache locality |
0.0 | 1 | 2002 | Creating and preserving locality of java applications at allocation and garbage collection times · OOPSLA 2002 |
Memory systems › cache
cache organization |
0.0 | 1 | 2002 | Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002 |
Emerging computing paradigms › unconventional computing
cellular computing |
0.0 | 1 | 2002 | Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002 |
Parallel and multicore computing › locality optimization
data locality optimization |
0.0 | 1 | 2002 | Creating and preserving locality of java applications at allocation and garbage collection times · OOPSLA 2002 |
Memory systems
memory hierarchy |
0.0 | 1 | 2002 | Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 2002 | Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002 |
Processor architecture and microarchitecture › multiprocessor architecture
synchronization hardware |
0.0 | 1 | 2002 | Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002 |
Integrated circuit design
system-on-chip |
0.0 | 1 | 2002 | An overview of the BlueGene/L Supercomputer · SC 2002 |
Compilers and program optimization
parallelizing compiler |
0.0 | 2 | 1996 | A Unified Framework for Optimizing Communication in Data-Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 An HPF Compiler for the IBM SP2 · SC 1995 |
Parallel and multicore computing
parallel programming models |
0.0 | 2 | 1996 | Static Analysis to Reduce Synchronization Costs in Data-Parallel Programs · POPL 1996 An HPF Compiler for the IBM SP2 · SC 1995 |
Compilers and program optimization › compiler optimization › redundancy elimination
array bounds check elimination |
0.0 | 1 | 2000 | From flop to megaflops: Java for technical computing · ACM Trans. Program. Lang. Syst. 2000 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 0.1connection graph · 0.1parallel structural analysis · 0.1hierarchical partitioning · 0.1MPI-IO · 0.1data flow analysis · 0.1performance analysis · 0.0time series analysis · 0.0rule-based classification · 0.0probabilistic networks · 0.0interprocedural analysis · 0.0bayesian network · 0.0simulation · 0.0prolific type analysis · 0.0locality-based graph traversal · 0.0static compilation · 0.0loop nest optimization · 0.0dynamic compilation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Modelling Using Deep LearningabstractMachine learning, and in particular, deep learning has emerged as an important tool for advancing science, in addition to its broad based impact on the world. This talk describes three research efforts that illustrate how deep learning can complement modeling and simulation to pursue scientific discoveries and to tackle societal problems. We begin by describing a flood forecasting initiative that has already led to hundreds of thousands of alerts being sent to people in India. It utilizes a new hydrologic model that has been built using an LSTM (long short-term memory) architecture and a physics based inundation model whose effectiveness has been enhanced using machine learning methods. We also describe how self supervised learning is being applied to study several interesting aspects of the organization of the human brain. The generated embeddings can be used to rapidly annotate new structures and develop new ways of clustering and categorizing brain structures based on purely data-driven criteria. Finally, we present a deep learning based modeling of human behavior in a specific game-based setting, which has very interesting implications if we are able to generalize that approach to broader settings. Manish Gupta 0002 |
ECMS | 1 |
| 2006 | High performance file I/O for the Blue Gene/L supercomputerabstractParallel I/O plays a crucial role for most data-intensive applications running on massively parallel systems like Blue Gene/L that provides the promise of delivering enormous computational capability. We designed and implemented a highly scalable parallel file I/O architecture for Blue Gene/L, which leverages the benefit of the hierarchical and functional partitioning design of the system software with separate computational and I/O cores. The architecture exploits the scalability aspect of GPFS (General Parallel File System) at the backend, while using MPI I/O as an interface between the application I/O and the file system. We demonstrate the impact of our high performance I/O solution for Blue Gene/L with a comprehensive evaluation that consists of a number of widely used parallel I/O benchmarks and I/O intensive applications. Our design and implementation is not only able to deliver at least one order of magnitude speed up in terms of I/O bandwidth for a real-scale application HOMME (achieving aggregate bandwidth of 1.8 GB/Sec and 2.3 GB/Sec for write and read accesses, respectively), but also supports high-level parallel I/O data interfaces such as parallel HDF5 and parallel NetCDF scaling up to a large number of processors. Hao Yu 0008, Ramendra K. Sahoo, C. Howson, Gheorghe Almási 0001, José G. Castaños, Manish Gupta 0002, José E. Moreira, Jeff Parker, Thomas Engelsiepen, Robert B. Ross, Rajeev Thakur, Robert Latham, William Gropp |
HPCA | 6 |
| 2006 | Gordon Bell finalists I - Large scale drop impact analysis of mobile phone using ADVC on Blue Gene/LabstractExisting commercial finite element analysis (FEA) codes do not exhibit the performance necessary for large scale analysis on parallel computer systems. In this paper, we demonstrate the performance characteristics of a commercial parallel structural analysis code, ADVC, on Blue Gene/L (BG/L). The numerical algorithm of ADVC is described, tuned, and optimized on BG/L, and then a large scale drop impact analysis of a mobile phone is performed. The model of the mobile phone is a nearly-full assembly that includes inner structures. The size of the model we have analyzed has 47 million nodal points and 142 million DOFs. This does not seem exceptionally large, but the dynamic impact analysis of a product model, with the contact condition on the entire surface of the outer case under this size, cannot be handled by other CAE systems. Our analysis is an unprecedented attempt in the electronics industry. It took only half a day, 12.1 hours, for the analysis of about 2.4 milliseconds. The floating point operation performance obtained has been 538 GFLOPS on 4096 node of BG/L. Hiroshi Akiba, Tomonobu Ohyama, Yoshinoir Shibata, Kiyoshi Yuyama, Yoshikazu Katai, Ryuichi Takeuchi, Takeshi Hoshino, Shinobu Yoshimura, Hirohisa Noguchi, Manish Gupta 0002, John A. Gunnels, Vernon Austel, Yogish Sabharwal, Rahul Garg 0001, Shoji Kato, Takashi Kawakami, Satoru Todokoro, Junko Ikeda |
SC | 10 |
| 2005 | Filtering Failure Logs for a BlueGene/L PrototypeabstractThe growing computational and storage needs of several scientific applications mandate the deployment of extreme-scale parallel machines, such as IBM's BlueGene/L, which can accommodate as many as 128K processors. In this paper, we present our experiences in collecting and filtering error event logs from a 8192 processor BlueGene/L prototype at IBM Rochester, which is currently ranked #8 in the Top-500 list. We analyze the logs collected from this machine over a period of 84 days starting from August 26, 2004. We perform a three-step filtering algorithm on these logs: extracting and categorizing failure events; temporal filtering to remove duplicate reports from the same location; and finally coalescing failure reports of the same error across different locations. Using this approach, we can substantially compress these logs, removing over 99.96% of the 828,387 original entries, and more accurately portray the failure occurrences on this system. Yinglung Liang, Yanyong Zhang, Anand Sivasubramaniam, Ramendra K. Sahoo, José E. Moreira, Manish Gupta 0002 |
DSN | 6 |
| 2005 | Probabilistic QoS Guarantees for Supercomputing SystemsabstractSupercomputing systems must be able to reliably and efficiently complete their assigned workloads, even in the presence of failures. This paper proposes a system that allows the system and users to negotiate a mutually desirable risk strategy; in order to accomplish this, the system makes probabilistic guarantees on quality of service (QoS), of the form, "Job j can be completed by deadline d with probability p". In order to make such guarantees, the system uses event prediction (forecasting) in conjunction with fault-aware job scheduling and cooperative checkpointing strategies. Using job logs and failure traces from actual high performance computing systems, we employ trace-based simulations to assess the effects of the prediction accuracy (a) and user risk strategy (U) on a variety of performance metrics. Compared to a system that does not use event prediction, a high forecasting accuracy resulted in QoS and utilization improvements of as much as 6%, along with an 89% reduction in the amount of lost work. Therefore, our results show that a system that makes probabilistic QoS guarantees using a market-based scheduling approach can increase both system performance and reliability. Adam J. Oliner, Larry Rudolph, Ramendra K. Sahoo, José E. Moreira, Manish Gupta 0002 |
DSN | 5 |
| 2005 | Early Experience with Scientific Applications on the Blue Gene/L Supercomputer
Gheorghe Almási 0001, Gyan Bhanot, Dong Chen 0005, Maria Eleftheriou, Blake G. Fitch, Alan Gara, Robert S. Germain, John A. Gunnels, Manish Gupta 0002, Philip Heidelberger, Michael Pitman, Aleksandr Rayshubskiy, James C. Sexton, Frank Suits, Pavlos Vranas, Robert Walkup, T. J. Christopher Ward, Yuriy Zhestkov, Alessandro Curioni, Wanda Andreoni, Charles Archer, José E. Moreira, Richard Loft, Henry M. Tufo, Theron Voran, Katherine Riley |
Euro-Par | 9 |
| 2005 | Scaling physics and material science applications on a massively parallel Blue Gene/L systemabstractBlue Gene/L represents a new way to build supercomputers, using a large number of low power processors, together with multiple integrated interconnection networks. Whether real applications can scale to tens of thousands of processors (on a machine like Blue Gene/L) has been an open question. In this paper, we describe early experience with several physics and material science applications on a 32,768 node Blue Gene/L system, which was installed recently at the Lawrence Livermore National Laboratory. Our study shows some problems in the applications and in the current software implementation, but overall, excellent scaling of these applications to 32K nodes on the current Blue Gene/L system. While there is clearly room for improvement, these results represent the first proof point that MPI applications can effectively scale to over ten thousand processors. They also validate the scalability of the hardware and software architecture of Blue Gene/L. Gheorghe Almási 0001, Gyan Bhanot, Alan Gara, Manish Gupta 0002, James C. Sexton, Robert Walkup, Vasily V. Bulatov, Andrew W. Cook, Bronis R. de Supinski, James N. Glosli, Jeffrey A. Greenough, François Gygi, Alison Kubota, Steve Louis, Thomas E. Spelce, Frederick H. Streitz, Peter L. Williams, Robert K. Yates, Charles Archer, José E. Moreira, Charles A. Rendleman |
ICS | 4 |
| 2004 | Finding and Removing Performance Bottlenecks in Large Systems
Glenn Ammons, Jong-Deok Choi, Manish Gupta 0002, Nikhil Swamy |
ECOOP | 3 |
| 2004 | Fault-Aware Job Scheduling for BlueGene/L SystemsabstractSummary form only given. Large-scale systems like BlueGene/L are susceptible to a number of software and hardware failures that can affect system performance. We evaluate the effectiveness of a previously developed job scheduling algorithm for BlueGene/L in the presence of faults. We have developed two new job-scheduling algorithms considering failures while scheduling the jobs. We have also evaluated the impact of these algorithms on average bounded slowdown, average response time and system utilization, considering different levels of proactive failure prediction and prevention techniques reported in the literature. Our simulation studies show that the use of these new algorithms with even trivial fault prediction confidence or accuracy levels (as low as 10%) can significantly improve the performance of the BlueGene/L system. Adam J. Oliner, Ramendra K. Sahoo, José E. Moreira, Manish Gupta 0002, Anand Sivasubramaniam |
IPDPS | 4 |
| 2004 | Whole-Stack Analysis and Optimization of Commercial Workloads on Server Systems
C. Richard Attanasio, Jong-Deok Choi, Niteesh Dubey, Kattamuri Ekanadham, Manish Gupta 0002, Tatsushi Inagaki, Kazuaki Ishizaki, Joefon Jann, Robert D. Johnson, Toshio Nakatani, Pratap Pattnaik, Mauricio J. Serrano, Stephen E. Smith, Ian M. Steiner, Yefim Shuf |
NPC | 5 |
| 2004 | Unlocking the Performance of the BlueGene/L SupercomputerabstractThe BlueGene/L supercomputer is expected to deliver new levels of application performance by providing a combination of good single-node computational performance and high scalability. To achieve good single-node performance, the BlueGene/L design includes a special dual floating-point unit on each processor and the ability to use two processors per node. BlueGene/L also includes both a torus and a tree network to achieve high scalability. We demonstrate how benchmarks and applications can take advantage of these architectural features to get the most out of BlueGene/L. Gheorghe Almási 0001, Siddhartha Chatterjee, Alan Gara, John A. Gunnels, Manish Gupta 0002, Amy Henning, José E. Moreira, Robert Walkup |
SC | 5 |
| 2003 | Critical event prediction for proactive management in large-scale computer clustersabstractAs the complexity of distributed computing systems increases, systems management tasks require significantly higher levels of automation; examples include diagnosis and prediction based on real-time streams of computer events, setting alarms, and performing continuous monitoring. The core of autonomic computing, a recently proposed initiative towards next-generation IT-systems capable of 'self-healing', is the ability to analyze data in real-time and to predict potential problems. The goal is to avoid catastrophic failures through prompt execution of remedial actions.This paper describes an attempt to build a proactive prediction and control system for large clusters. We collected event logs containing various system reliability, availability and serviceability (RAS) events, and system activity reports (SARs) from a 350-node cluster system for a period of one year. The 'raw' system health measurements contain a great deal of redundant event data, which is either repetitive in nature or misaligned with respect to time. We applied a filtering technique and modeled the data into a set of primary and derived variables. These variables used probabilistic networks for establishing event correlations through prediction algorithms. We also evaluated the role of time-series methods, rule-based classification algorithms and Bayesian network models in event prediction.Based on historical data, our results suggest that it is feasible to predict system performance parameters (SARs) with a high degree of accuracy using time-series models. Rule-based classification techniques can be used to extract machine-event signatures to predict critical events with up to 70% accuracy. Ramendra K. Sahoo, Adam J. Oliner, Irina Rish, Manish Gupta 0002, José E. Moreira, Sheng Ma, Ricardo Vilalta, Anand Sivasubramaniam |
KDD | 4 |
| 2003 | Enabling Dual-Core Mode in BlueGene/L: Challenges and SolutionsabstractBlueGene/L is a massively parallel computer system with 65536 dual-processor compute nodes. The peak performance of BlueGene/L is in excess of 360 TFLOP/s if both processor cores in a node are used for computation. The main challenge of deploying this dual-core mode of operation is that the L1 caches in each core are not hardware coherent. This forces a software-based approach to cache coherence and guides our design of a programming model for dual-core mode. We describe the design, implementation, and performance evaluation of system software for enabling the use of dual-core mode on BlueGene/L. Our preliminary performance results show that our approach to dual-core mode is effective for key numerical kernels. George S. Almási, Leonardo R. Bachega, Siddhartha Chatterjee, Manish Gupta 0002, Derek Lieber, Xavier Martorell, José E. Moreira |
SBAC-PAD | 4 |
| 2003 | Supporting multidimensional arrays in JavaabstractAbstract The lack of direct support for multidimensional arrays in JavaTM has been recognized as a major deficiency in the language's applicability to numerical computing. It has been shown that, when augmented with multidimensional arrays, Java can achieve very high‐performance for numerical computing through the use of compiler techniques and efficient implementations of aggregate array operations. Three approaches have been discussed in the literature for extending Java with support for multidimensional arrays: class libraries that implement these structures; extending the Java language with new syntactic constructs for multidimensional arrays that are directly translated to bytecode; and relying on the Java Virtual Machine to recognize those arrays of arrays that are being used to simulate multidimensional arrays. This paper presents a balanced presentation of the pros and cons of each technique in the areas of functionality, language and virtual machine impact, implementation effort, and effect on performance. We show that the best choice depends on the relative importance attached to the different metrics, and thereby provide a common ground for a rational discussion and comparison of the techniques. Copyright © 2003 John Wiley & Sons, Ltd. José E. Moreira, Samuel P. Midkiff, Manish Gupta 0002 |
Concurr. Comput. Pract. Exp. | 3 |
| 2003 | Stack allocation and synchronization optimizations for Java using escape analysisabstractThis article presents an escape analysis framework for Java to determine (1) if an object is not reachable after its method of creation returns, allowing the object to be allocated on the stack, and (2) if an object is reachable only from a single thread during its lifetime, allowing unnecessary synchronization operations on that object to be removed. We introduce a new program abstraction for escape analysis, the connection graph , that is used to establish reachability relationships between objects and object references. We show that the connection graph can be succinctly summarized for each method such that the same summary information may be used in different calling contexts without introducing imprecision into the analysis. We present an interprocedural algorithm that uses the above property to efficiently compute the connection graph and identify the nonescaping objects for methods and threads. The experimental results, from a prototype implementation of our framework in the IBM High Performance Compiler for Java, are very promising. The percentage of objects that may be allocated on the stack exceeds 70% of all dynamically created objects in the user code in three out of the ten benchmarks (with a median of 19%); 11% to 92% of all mutex lock operations are eliminated in those 10 programs (with a median of 51%), and the overall execution time reduction ranges from 2% to 23% (with a median of 7%) on a 333-MHz PowerPC workstation with 512 MB memory. Jong-Deok Choi, Manish Gupta 0002, Mauricio J. Serrano, Vugranam C. Sreedhar, Samuel P. Midkiff |
ACM Trans. Program. Lang. Syst. | 2 |
| 2002 | Blue Gene/L, a System-On-A-ChipabstractSummary form only given. Large powerful networks coupled to state-of-the-art processors have traditionally dominated supercomputing. As technology advances, this approach is likely to be challenged by a more cost-effective System-On-A-Chip approach, with higher levels of system integration. The scalability of applications to architectures with tens to hundreds of thousands of processors is critical to the success of this approach. Significant progress has been made in mapping numerous compute-intensive applications, many of them grand challenges, to parallel architectures. Applications hoping to efficiently execute on future supercomputers of any architecture must be coded in a manner consistent with an enormous degree of parallelism. The BG/L program is developing a peak nominal 180 TFLOPS (360 TFLOPS for some applications) supercomputer to serve a broad range of science applications. BG/L generalizes QCDOC, the first System-On-A-Chip supercomputer that is expected in 2003. BG/L consists of 65,536 nodes, and contains five integrated networks: a 3D torus, a combining tree, a Gb Ethernet network, barrier/global interrupt network and JTAG. George S. Almási, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, Alina Deutsch, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, Blake G. Fitch, Joseph Gagliano, Alan Gara, Robert S. Germain, Mark Giampapa, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, T. Jamal-Eddine, Gerard V. Kopcsay, Alphonso P. Lanzetta, Derek Lieber, M. Lu, Mark P. Mendell, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Rick A. Rand, Richard D. Regan, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, Sarabjeet Singh, Peilin Song, Burkhard D. Steinmacher-Burow, Karin Strauss, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, J. Brown, Thomas A. Liebsch, A. Schram, G. Ulsh |
CLUSTER | 27 |
| 2002 | Evaluation of a Multithreaded Architecture for Cellular ComputingabstractCyclops is a new architecture for high-performance parallel computers that is being developed at the IBM T. J. Watson Research Center. The basic cell of this architecture is a single-chip SMP (symmetric multiprocessor) system with multiple threads of execution, embedded memory and integrated communications hardware. Massive intra-chip parallelism is used to tolerate memory and functional unit latencies. Large systems with thousands of chips can be built by replicating this basic cell in a regular pattern. In this paper, we describe the Cyclops architecture and evaluate two of its new hardware features: a memory hierarchy with a flexible cache organization and fast barrier hardware. Our experiments with the STREAM benchmark show that a particular design can achieve a sustainable memory bandwidth of 40 GB/s, equal to the peak hardware bandwidth and similar to the performance of a 128-processor SGI Origin 3800. For small vectors, we have observed in-cache bandwidth above 80 GB/s. We also show that the fast barrier hardware can improve the performance of the Splash-2 FFT kernel by up to 10%. Our results demonstrate that the Cyclops approach of integrating a large number of simple processing elements and multiple memory banks in the same chip is an effective alternative for designing high-performance systems. Calin Cascaval, José G. Castaños, Luis Ceze, Monty Denneau, Manish Gupta 0002, Derek Lieber, José E. Moreira, Karin Strauss, Henry S. Warren Jr. |
HPCA | 5 |
| 2002 | Creating and preserving locality of java applications at allocation and garbage collection timesabstractThe growing gap between processor and memory speeds is motivating the need for optimization strategies that improve data locality. A major challenge is to devise techniques suitable for pointer-intensive applications. This paper presents two techniques aimed at improving the memory behavior of pointer-intensive applications with dynamic memory allocation, such as those written in Java. First, we present an allocation time object placement technique based on the recently introduced notion of prolific (frequently instantiated) types. We attempt to co-locate, at allocation time, objects of prolific types that are connected via object references. Then, we present a novel locality based graph traversal technique. The benefits of this technique, when applied to garbage collection (GC), are twofold: (i) it improves the performance of GC due to better locality during a heap traversal and (ii) it restructures surviving objects in a way that enhances locality. On multiprocessors, this technique can further reduce overhead due to synchronization and false sharing. The experimental results, on a well-known suite of Java benchmarks (SPECjvm98 [26], SPECjbb2000 [27], and jOlden [4]), from an implementation of these techniques in the Jikes RVM [1], are very encouraging. The object co-allocation technique improves application performance by up to 21% (10% on average) in the Jikes RVM configured with a non-copying mark-and-sweep collector. The locality-based traversal technique reduces GC times by up to 20% (10% on average) and improves the performance of applications by up to 14% (6% on average) in the Jikes RVM configured with a copying semi-space collector. Both techniques combined can improve application performance by up to 22% (10% on average) in the Jikes RVM configured with a non-copying mark-and-sweep collector. Yefim Shuf, Manish Gupta 0002, Hubertus Franke, Andrew W. Appel, Jaswinder Pal Singh |
OOPSLA | 2 |
| 2002 | Exploiting prolific types for memory management and optimizationsabstractIn this paper, we introduce the notion of prolific and non-prolific types, based on the number of instantiated objects of those types. We demonstrate that distinguishing between these types enables a new class of techniques for memory management and data locality, and facilitates the deployment of known techniques. Specifically, we first present a new type-based approach to garbage collection that has similar attributes but lower cost than generational collection. Then we describe the short type pointer technique for reducing memory requirements of objects (data) used by the program. We also discuss techniques to facilitate the recycling of prolific objects and to simplify object co-allocation decisions.We evaluate the first two techniques on a standard set of Java benchmarks (SPECjvm98 and SPECjbb2000). An implementation of the type-based collector in the Jalapeño VM shows improved pause times, elimination of unnecessary write barriers, and reduction in garbage collection time (compared to the analogous generational collector) by up to 15%. A study to evaluate the benefits of the short-type pointer technique shows a potential reduction in the heap space requirements of programs by up to 16%. Yefim Shuf, Manish Gupta 0002, Rajesh Bordawekar, Jaswinder Pal Singh |
POPL | 2 |
| 2002 | An overview of the BlueGene/L SupercomputerabstractThis paper gives an overview of the BlueGene/L Supercomputer. This is a jointly funded research partnership between IBM and the Lawrence Livermore National Laboratory as part of the United States Department of Energy ASCI Advanced Architecture Research Program. Application performance and scaling studies have recently been initiated with partners at a number of academic and government institutions,including the San Diego Supercomputer Center and the California Institute of Technology. This massively parallel system of 65,536 nodes is based on a new architecture that exploits system-on-a-chip technology to deliver target peak processing power of 360 teraFLOPS (trillion floating-point operations per second). The machine is scheduled to be operational in the 2004-2005 time frame, at price/performance and power consumption/performance targets unobtainable with conventional architectures. Narasimha R. Adiga, Gheorghe Almási 0001, George S. Almási, Yariv Aridor, Rajkishore Barik, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Waiman Chan, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, K. M. Desai, Alina Deutsch, Tamar Domany, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, C. Christopher Erway, J. Esch, Blake G. Fitch, Joseph Gagliano, Alan Gara, Rahul Garg 0001, Robert S. Germain, Mark Giampapa, Balaji Gopalsamy, John A. Gunnels, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, R. D. Jackson, T. Jamal-Eddine, Gerard V. Kopcsay, Elie Krevat, Manish P. Kurhekar, Alphonso P. Lanzetta, Derek Lieber, L. K. Liu, M. Lu, Mark P. Mendell, A. Misra, Yosef Moatti, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Adam J. Oliner, Vinayaka Pandit, R. B. Pudota, Rick A. Rand, Richard D. Regan, Bradley Rubin, Albert E. Ruehli, Silvius Vasile Rus, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, M. Sharma, Edi Shmueli, Sarabjeet Singh, Peilin Song, Vijay Srinivasan, Burkhard D. Steinmacher-Burow, Karin Strauss, Christopher W. Surovic, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Arun R. Umamaheshwaran, P. Verma, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, W. Barrett, C. Engel, B. Drehmel, B. Hilgart, D. Hill, F. Kasemkhani, David J. Krolak, Chun-Tao Li 0001, Thomas A. Liebsch, James A. Marcella, A. Muff, A. Okomo, M. Rouse, A. Schram, M. Tubbs, G. Ulsh, Charles D. Wait, J. Wittrup, Myung Bae, Kenneth A. Dockser, Lynn Kissel, Mark K. Seager, Jeffrey S. Vetter, K. Yates |
SC | 39 |
| 2001 | A framework for efficient reuse of binary code in JavaabstractThis paper presents a compilation framework that enables efficient sharing of executable code across distinct Java Virtual Machine (JVM) instances. High-performance JVMs rely on run-time compilation, since static compilation cannot handle many dynamic features of Java. These JVMs suffer from large memory footprints and high startup costs, which are serious problems for embedded devices (such as hand held personal digital assistants and cellular phones) and scalable servers. A recently proposed approach called quasi-static compilation overcomes these difficulties by reusing precompiled binary images after performing validation checks and stitching on them (i.e., adapting them to a new execution context), falling back to interpretation or dynamic compilation whenever necessary. However, the requirement in our previous design to duplicate and modify the executable binary image for stitching is a major drawback when targeting embedded systems and scalable servers. In this paper, we describe a new approach that allows stitching to be done on an indirection table, leaving the executable code unmodified and therefore writable to readonly memory. On embedded devices, this saves precious space in writable memory. On scalable servers, this allows a single image of the executable to be shared among multiple JVMs, thus improving scalability. Furthermore, we describe a novel technique for dynamically linking classes that uses traps to detect when a class should be linked and initialized. Like back-patching, the technique allows all accesses after the first to proceed at full speed, but unlike back-patching, it avoids the modification of running code. We have implemented this approach in the Quicksilver quasi-static com- Pramod G. Joisha, Samuel P. Midkiff, Mauricio J. Serrano, Manish Gupta 0002 |
ICS | 4 |
| 2000 | Optimizing Java Programs in the Presence of Exceptions
Manish Gupta 0002, Jong-Deok Choi, Michael Hind |
ECOOP | 1 |
| 2000 | Automatic loop transformations and parallelization for JavaabstractFrom a software engineering perspective, the Java programming language provides an attractive platform for writing numerically intensive applications. A major drawback hampering its widespread adoption in this domain has been its poor performance on numerical codes. This paper describes a prototype Java compiler which demonstrates that it is possible to achieve performance levels approaching those of current state-of-the-art C, C++ and Fortran compilers on numerical codes. We describe a new transformation called alias versioning that takes advantage of the simplicity of pointers in Java. This transformation, combined with other techniques that we have developed, enables the compiler to perform high order loop transformations (for better data locality) and parallelization completely automatically. We believe that our compiler is the first to have such capabilities of optimizing numerical Java codes. We achieve, with Java, between 80 and 100% of the performance of highly optimized Fortran code in a variety of benchmarks. Furthermore, the automatic parallelization achieves speedups of up to 3.8 on four processors. Combining this compiler technology with packages containing the features expected by programmers of numerical applications would enable Java to become a serious contender for implementing new numerical applications. Pedro V. Artigas, Manish Gupta 0002, Samuel P. Midkiff, José E. Moreira |
ICS | 2 |
| 2000 | Quicksilver: a quasi-static compiler for JavaabstractThis paper presents the design and implementation of the Quicksilver1 quasi-static compiler for Java. Quasi-static compilation is a new approach that combines the benefits of static and dynamic compilation, while maintaining compliance with the Java standard, including support of its dynamic features. A quasi-static compiler relies on the generation and reuse of persistent code images to reduce the overhead of compilation during program execution, and to provide identical, testable and reliable binaries over different program executions. At runtime, the quasi-static compiler adapts pre-compiled binaries to the current JVM instance, and uses dynamic compilation of the code when necessary to support dynamic Java features. Our system allows interprocedural program optimizations to be performed while maintaining binary compatibility. Experimental data obtained using a preliminary implementation of a quasi-static compiler in the Jalapeño JVM clearly demonstrates the benefits of our approach: we achieve a runtime compilation cost comparable to that of baseline (fast, non-optimizing) compilation, and deliver the runtime program performance of the highest optimization level supported by the Jalapeño optimizing compiler. For the SPECjvm98 benchmark suite, we obtain a factor of 104 to 158 reduction in the runtime compilation overhead relative to the Jalapeño optimizing compiler. Relative to the better of the baseline and the optimizing Jalapeño compilers, the overall performance (taking into account both runtime compilation and execution costs) is increased by 9.2% to 91.4% for the SPECjvm98 benchmarks with size 100, and by 54% to 356% for the (shorter running) SPECjvm98 benchmarks with size 10. Mauricio J. Serrano, Rajesh Bordawekar, Samuel P. Midkiff, Manish Gupta 0002 |
OOPSLA | 4 |
| 2000 | From flop to megaflops: Java for technical computingabstractAlthough there has been some experimentation with Java as a language for numerically intensive computing, there is a perception by many that the language is unsuited for such work because of performance deficiencies. In this article we show how optimizing array bounds checks and null pointer checks creates loop nests on which aggressive optimizations can be used. Applying these optimizations by hand to a simple matrix-multiply test case leads to Java-compliant programs whose performance is in excess of 500 Mflops on a four-processor 332MHz RS/6000 model F50 computer. We also report in this article the effect that various optimizations have on the performance of six floating-point-intensive benchmarks. Through these optimizations we have been able to achieve with Java at least 80% of the peak Fortran performance on the same benchmarks. Since all of these optimizations can be automated, we conclude that Java will soon be a serious contender for numerically intensive computing. José E. Moreira, Samuel P. Midkiff, Manish Gupta 0002 |
ACM Trans. Program. Lang. Syst. | 3 |
| 1999 | Escape Analysis for JavaabstractThis paper presents a simple and efficient data flow algorithm for escape analysis of objects in Java programs to determine (i) if an object can be allocated on the stack; (ii) if an object is accessed only by a single thread during its lifetime, so that synchronization operations on that object can be removed. We introduce a new program abstraction for escape analysis, the connection graph, that is used to establish reachability relationships between objects and object references. We show that the connection graph can be summarized for each method such that the same summary information may be used effectively in different calling contexts. We present an interprocedural algorithm that uses the above property to efficiently compute the connection graph and identify the non-escaping objects for methods and threads. The experimental results, from a prototype implementation of our framework in the IBM High Performance Compiler for Java, are very promising. The percentage of objects that may be allocated on the stack exceeds 70% of all dynamically created objects in three out of the ten benchmarks (with a median of 19%), 11% to 92% of all lock operations are eliminated in those ten programs (with a median of 51%), and the overall execution time reduction ranges from 2% to 23% (with a median of 7%) on a 333 MHz PowerPC workstation with 128 MB memory. Jong-Deok Choi, Manish Gupta 0002, Mauricio J. Serrano, Vugranam C. Sreedhar, Samuel P. Midkiff |
OOPSLA | 2 |
| 1999 | High Performance Computing with the Array Package for Java: A Case Study using Data MiningabstractThis paper discusses several techniques used in developing a parallel, production quality data mining application in Java.We started by developing three sequential versions of a product recommendation data mining application: (i) a Fortran 90 version used as a performance reference, (ii) a plain Java implementation that only uses the primitive array structures from the language, and (iii) a baseline Java implementation that uses our Array package for Java.This Array package provides parallelism at the level of individual Array and BLAS operations.Using this Array package, we also developed two parallel Java versions of the data mining application: one that relies entirely on the implicit parallelism provided by the Array package, and another that is explicitly parallel at the application level.We discuss the design of the Array package, as well as the design of the data mining application.We compare the trade-offs between performance and the abstraction level the different Java versions present to the application programmer.Our studies show that, although a plain Java implementation performs poorly, the Java implementation with the Array package is quite competitive in performance with Fortran.We achieve a single processor performance of 109 Mflops, or 91% of Fortran performance, on a 332 MHz PowerPC 604e processor.Both the implicitly and explicitly parallel forms of our Java implementations also parallelize well.On an SMP with four of those PowerPC processors, the implicitly parallel form achieves 290 Mflops with no effort from the application programmer, while the explicitly parallel form achieves 340 Mflops.1 José E. Moreira, Samuel P. Midkiff, Manish Gupta 0002, Rick Lawrence |
SC | 3 |
| 1996 | Global Communication Analysis and OptimizationabstractReducing communication cost is crucial to achieving good performance on scalable parallel machines. This paper presents a new compiler algorithm for global analysis and optimization of communication in data-parallel programs. Our algorithm is distinct from existing approaches in that rather than handling loop-nests and array references one by one, it considers all communication in a procedure and their interactions under different placements before making a final decision on the placement of any communication. It exploits the flexibility resulting from this advanced analysis to eliminate redundancy, reduce the number of messages, and reduce contention for cache and communication buffers, all in a unified framework. In contrast, single loop-nest analysis often retains redundant communication, and more aggressive dataflow analysis on array sections can generate too many messages or cache and buffer contention. The algorithm has been implemented in the IBM pHPF compiler for High Performance Fortran. During compilation, the number of messages per processor goes down by as much as a factor of nine for some HPF programs. We present performance results for the IBM SP2 and a network of Sparc workstations (NOW) connected by a Myrinet switch. In many cases, the communication cost is reduced by a factor of two. Soumen Chakrabarti, Manish Gupta 0002, Jong-Deok Choi |
PLDI | 2 |
| 1996 | Static Analysis to Reduce Synchronization Costs in Data-Parallel ProgramsabstractFor a program with sufficient parallelism, reducing synchronization costs is one of the most important objectives for achieving efficient execution on any parallel machine. This paper presents a novel methodology for reducing synchronization costs of programs compiled for SPMD execution. This methodology combines data flow analysis with communication analysis to determine the ordering between production and consumption of data on different processors, which helps in identifying redundant synchronization. The resulting framework is more powerful than any that have been previously presented, as it provides the first algorithm that can eliminate synchronization messages even from computations that need communication. We show that several commonly occurring computation patterns such as reductions and stencil computations with reciprocal producer-consumer relationship between processors lend themselves well to this optimization, an observation that is confirmed by an examination of some HPF benchmark programs. Our framework also recognizes situations where the synchronization needs for multiple data transfers can be satisfied by a single synchronization message. This analysis, while applicable to all shared memory machines as well, is especially useful for those with a flexible cache-coherence protocol, as it identifies efficient ways of moving data directly from producers to consumers, often without any extra synchronization. Manish Gupta 0002, Edith Schonberg |
POPL | 1 |
| 1996 | A Unified Framework for Optimizing Communication in Data-Parallel ProgramsabstractThis paper presents a framework, based on global array data-flow analysis, to reduce communication costs in a program being compiled for a distributed memory machine. We introduce available section descriptor, a novel representation of communication involving array sections. This representation allows us to apply techniques for partial redundancy elimination to obtain powerful communication optimizations. With a single framework, we are able to capture optimizations like (1) vectorizing communication, (2) eliminating communication that is redundant on any control flow path, (3) reducing the amount of data being communicated, (4) reducing the number of processors to which data must be communicated, and (5) moving communication earlier to hide latency, and to subsume previous communication. We show that the bidirectional problem of eliminating partial redundancies can be decomposed into simpler unidirectional problems even in the context of an array section representation, which makes the analysis procedure more efficient. We present results from a preliminary implementation of this framework, which are extremely encouraging, and demonstrate the effectiveness of this analysis in improving the performance of programs. Manish Gupta 0002, Edith Schonberg, Harini Srinivasan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1995 | An HPF Compiler for the IBM SP2abstractWe describe pHPF, an research prototype HPF compiler for the IBM SP series parallel machines. The compiler accepts as input Fortran 90 and Fortran 77 programs, augmented with HPF directives; sequential loops are automatically parallelized. The compiler supports symbolic analysis of expressions. This allows parameters such as the number of processors to be unknown at compile-time without significantly affecting performance. Communication schedules and computation guards are generated in a parameterized form at compile-time. Several novel optimizations and improved versions of well-known optimizations have been implemented in pHPF to exploit parallelism and reduce communication costs. These optimizations include elimination of redundant communication using data-availability analysis; using collective communication; new techniques for mapping scalar variables; coarse-grain wavefronting; and communication reduction in multi-dimensional shift communications. We present experimenta... Manish Gupta 0002, Samuel P. Midkiff, Edith Schonberg, Ven Seshadri, David Shields, Ko-Yang Wang, Wai-Mee Ching, Ton Anh Ngo |
SC | 1 |
| 1991 | Effects of Program Parallelization and Stripmining Transformation on Cache Performance in a Multiprocessor
Manish Gupta 0002, David A. Padua |
ICPP (1) | 1 |