Toshio Suganuma

dblp:35/1392 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 4 first-authorSystems, architecture and hardware · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Runtime systems and virtual machines · 48% Compilers and program optimization · 42% Program analysis · 10%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.252006
A region-based compilation technique for dynamic compilers · ACM Trans. Program. Lang. Syst. 2006
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
A region-based compilation technique for a Java just-in-time compiler · PLDI 2003
Compilers and program optimization › interprocedural optimization
inlining
0.232006
A region-based compilation technique for dynamic compilers · ACM Trans. Program. Lang. Syst. 2006
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Runtime systems and virtual machines › dynamic compilation › just-in-time compilation
region-based compilation
0.122006
A region-based compilation technique for dynamic compilers · ACM Trans. Program. Lang. Syst. 2006
A region-based compilation technique for a Java just-in-time compiler · PLDI 2003
Compilers and program optimization
dynamic optimization
0.122005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
A Dynamic Optimization Framework for a Java Just-In-Time Compiler · OOPSLA 2001
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.122005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
A Dynamic Optimization Framework for a Java Just-In-Time Compiler · OOPSLA 2001
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.122005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
A Dynamic Optimization Framework for a Java Just-In-Time Compiler · OOPSLA 2001
Runtime systems and virtual machines › dynamic compilation
on-stack replacement
0.122006
A region-based compilation technique for dynamic compilers · ACM Trans. Program. Lang. Syst. 2006
A region-based compilation technique for a Java just-in-time compiler · PLDI 2003
Program analysis
data flow analysis
0.122006
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
A region-based compilation technique for dynamic compilers · ACM Trans. Program. Lang. Syst. 2006
Compilers and program optimization › dynamic optimization
adaptive compilation
0.112005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
Compilers and program optimization
program specialization
0.112005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005
Program analysis › dynamic analysis › profiling
value profiling
0.012001
A Dynamic Optimization Framework for a Java Just-In-Time Compiler · OOPSLA 2001
Program analysis › dynamic analysis
profiling
0.012005
Design and evaluation of dynamic optimizations for a Java just-in-time compiler · ACM Trans. Program. Lang. Syst. 2005

Methods — techniques the papers use, named apart from their topics

static heuristics · 0.1dynamic profiling · 0.1instrumentation · 0.1value profiling · 0.1on-stack replacement · 0.1partial redundancy elimination · 0.0exception check elimination · 0.0sampling profiler · 0.0
YearPublicationVenuePosition
2011 Distributed and fault-tolerant execution framework for transaction processing
abstract
There is a growing need for efficient distributed computing for transaction processing. One of the key requirements for runtime systems in distributed environments is fault tolerance. Such a system needs to preserve the data consistency at transaction boundaries so as to resume the ongoing tasks from checkpoints with consistent data for any component failure. Another key requirement is that the system needs to be lightweight enough in normal execution to provide scalable performance. This paper presents the design and implementation of a new fault tolerant execution framework that addresses both of these requirements. We replicate each partition of the distributed persistent data on three nodes (triplet) with two different types of backups, one using warm replication and the other using cold replication. For node failures, the system is automatically recoverable unless all three nodes in any triplet fail at the same time. The system tolerates simultaneous two-node failures in any triplet most of the cases. We obtained a new trade-off in that 43% performance improvements can be achieved by slightly compromising the system availability.
Toshio Suganuma, Akira Koseki, Kazuaki Ishizaki, Yohei Ueda, Ken Mizuno, Daniel Silva 0001, Hideaki Komatsu, Toshio Nakatani
SYSTOR1
2010 Parallel programming framework for large batch transaction processing on scale-out systems
abstract
A scale-out system is a cluster of commodity machines, and offers a good platform to support steadily increasing workloads that process growing data sets. Sharding [4] is a method of partitioning data and processing a computation on a scale-out system. In a database system, a large table can be partitioned into small tables so each node can process its part of the computation. The sharding approach in a large batch transaction processing, which is important in financial area, presents two hard problems to programmers. Programmers have to write complex code (1) to transfer the input data so as to align the computations with the data partitions, and (2) to manage the distributed transactions. This paper presents a new parallel programming framework that makes parallel transactional programming easier by specifying transaction scopes and partitioners to simplify the code. Transaction scopes include series of subtransactions, each of which performs local operations. The system manages the distributed transactions automatically. A partitioner represents how the computation should be decomposed and aligned with the data partitions to avoid remote database accesses. Between paired of subtransactions, the system handles the data shuffling across the network. We implemented our parallel programming framework as a new Java class library. We hide all of the complex details of data transfer and distributed transaction management in the library. Our programming framework can eliminate almost 66% of the lines of code compared to a current programming approach without programming framework support. We also confirmed good scalability, with a scaling factor of 20.6 on 24 nodes using our modified batch program for the TPC-C benchmark.
Kazuaki Ishizaki, Ken Mizuno, Toshio Suganuma, Daniel Silva 0001, Akira Koseki, Hideaki Komatsu, Yohei Ueda, Toshio Nakatani
SYSTOR3
2006 A region-based compilation technique for dynamic compilers
abstract
Method inlining and data flow analysis are two major optimization components for effective program transformations, but they often suffer from the existence of rarely or never executed code contained in the target method. One major problem lies in the assumption that the compilation unit is partitioned at method boundaries. This article describes the design and implementation of a region-based compilation technique in our dynamic optimization framework, in which the compiled regions are selected as code portions without rarely executed code. The key parts of this technique are the region selection, partial inlining, and region exit handling. For region selection, we employ both static heuristics and dynamic profiles to identify and eliminate rare sections of code. The region selection process and method inlining decisions are interwoven, so that method inlining exposes other targets for region selection, while the region selection in the inline target conserves the inlining budget, allowing more method inlining to be performed. The inlining process can be performed for parts of a method, not just for the entire body of the method. When the program attempts to exit from a region boundary, we trigger recompilation and then use on-stack replacement to continue the execution from the corresponding entry point in the recompiled code. We have implemented these techniques in our Java JIT compiler, and conducted a comprehensive evaluation. The experimental results show that our region-based compilation approach achieves approximately 4% performance improvement on average, while reducing the compilation overhead by 10% to 30%, in comparison to the traditional method-based compilation techniques.
Toshio Suganuma, Toshiaki Yasue, Toshio Nakatani
ACM Trans. Program. Lang. Syst.1
2005 Design and evaluation of dynamic optimizations for a Java just-in-time compiler
abstract
The high performance implementation of Java Virtual Machines (JVM) and Just-In-Time (JIT) compilers is directed toward employing a dynamic compilation system on the basis of online runtime profile information. The trade-off between the compilation overhead and performance benefit is a crucial issue for such a system. This article describes the design and implementation of a dynamic optimization framework in a production-level Java JIT compiler, together with two techniques for profile-directed optimizations: method inlining and code specialization. Our approach is to employ a mixed mode interpreter and a three-level optimizing compiler, supporting level-1 to level-3 optimizations, each of which has a different set of trade-offs between compilation overhead and execution speed. A lightweight sampling profiler operates continuously during the entire period while applications are running to monitor the programs' hot spots. Detailed information on runtime behavior can be collected by dynamically generating instrumentation code that is installed to and uninstalled from the specified recompilation target code. Value profiling with this instrumentation mechanism allows fully automatic profile-directed method inlining and code specialization to be performed on the basis of call site information or specific parameter values at the higher optimization levels. The experimental results show that our approach offers high performance and low compilation overhead in both program startup and steady state measurements in comparison to the previous systems. The two profile-directed optimization techniques contribute significant portions of the improvements.
Toshio Suganuma, Toshiaki Yasue, Motohiro Kawahito, Hideaki Komatsu, Toshio Nakatani
ACM Trans. Program. Lang. Syst.1
2003 Effectiveness of cross-platform optimizations for a java just-in-time compiler
abstract
This paper describes the system overview of our Java Just-In-Time (JIT) compiler, which is the basis for the latest production version of IBM Java JIT compiler that supports a diversity of processor architectures including both 32-bit and 64-bit modes, CISC, RISC, and VLIW architectures. In particular, we focus on the design and evaluation of the cross-platform optimizations that are common across different architectures. We studied the effectiveness of each optimization by selectively disabling it in our JIT compiler on three different platforms: IA-32, IA-64, and PowerPC. Our detailed measurements allowed us to rank the optimizations in terms of the greatest performance improvements with the smallest compilation times. The identified set includes method inlining only for tiny methods, exception check eliminations using forward dataflow analysis and partial redundancy elimination, scalar replacement for instance and class fields using dataflow analysis, optimizations for type inclusion checks, and the elimination of merge points in the control flow graphs. These optimizations can achieve 90% of the peak performance for two industry-standard benchmark programs on these platforms with only 34% of the compilation time compared to the case for using all of the optimizations.
Kazuaki Ishizaki, Mikio Takeuchi, Kiyokuni Kawachiya, Toshio Suganuma, Osamu Gohda, Tatsushi Inagaki, Akira Koseki, Kazunori Ogata, Motohiro Kawahito, Toshiaki Yasue, Takeshi Ogasawara, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
OOPSLA4
2003 A region-based compilation technique for a Java just-in-time compiler
abstract
Method inlining and data flow analysis are two major optimization components for effective program transformations, however they often suffer from the existence of rarely or never executed code contained in the target method. One major problem lies in the assumption that the compilation unit is partitioned at method boundaries. This paper describes the design and implementation of a region-based compilation technique in our dynamic compilation system, in which the compiled regions are selected as code portions without rarely executed code. The key part of this technique is the region selection, partial inlining, and region exit handling. For region selection, we employ both static heuristics and dynamic profiles to identify rare sections of code. The region selection process and method inlining decision are interwoven, so that method inlining exposes other targets for region selection, while the region selection in the inline target conserves the inlining budget, leading to more method inlining. Thus the inlining process can be performed for parts of a method, not for the entire body of the method. When the program attempts to exit from a region boundary, we trigger recompilation and then rely on on-stack replacement to continue the execution from the corresponding entry point in the recompiled code. We have implemented these techniques in our Java JIT compiler, and conducted a comprehensive evaluation. The experimental results show that the approach of region-based compilation achieves approximately 5% performance improvement on average, while reducing the compilation overhead by 20 to 30%, in comparison to the traditional function-based compilation techniques.
Toshio Suganuma, Toshiaki Yasue, Toshio Nakatani
PLDI1
2001 A Dynamic Optimization Framework for a Java Just-In-Time Compiler
abstract
The high performance implementation of Java Virtual Machines (JVM) and just-in-time (JIT) compilers is directed toward adaptive compilation optimizations on the basis of online runtime profile information. This paper describes the design and implementation of a dynamic optimization framework in a production-level Java JIT compiler. Our approach is to employ a mixed mode interpreter and a three level optimizing compiler, supporting quick, full, and special optimization, each of which has a different set of tradeoffs between compilation overhead and execution speed. a lightweight sampling profiler operates continuously during the entire program's exectuion. When necessary, detailed information on runtime behavior is collected by dynmiacally generating instrumentation code which can be installed to and uninstalled from the specified recompilation target code. Value profiling with this instrumentation mechanism allows fully automatic code specialization to be performed on the basis of specific parameter values or global data at the highest optimization level. The experimental results show that our approach offers high performance and a low code expansion ratio in both program startup and steady state measurements in comparison to the compile-only approach, and that the code specialization can also contribute modest performance improvement
Toshio Suganuma, Toshiaki Yasue, Motohiro Kawahito, Hideaki Komatsu, Toshio Nakatani
OOPSLA1
2000 Design, implementation, and evaluation of optimizations in a JavaTM Just-In-Time compiler
abstract
The Java language incurs a runtime overhead for exception checks and object accesses, which are executed without an interior pointer in order to ensure safety. It also requires type inclusion test, dynamic class loading, and dynamic method calls in order to ensure flexibility. A ‘Just-In-Time’ (JIT) compiler generates native code from Java byte code at runtime. It must improve the runtime performance without compromising the safety and flexibility of the Java language. We designed and implemented effective optimizations for a JIT compiler, such as exception check elimination, common subexpression elimination, simple type inclusion test, method inlining, and devirtualization of dynamic method call. We evaluate the performance benefits of these optimizations based on various statistics collected using SPECjvm98, its candidates, and two JavaSoft applications with byte code sizes ranging from 23 000 to 280 000 bytes. Each optimization contributes to an improvement in the performance of the programs. Copyright © 2000 John Wiley & Sons, Ltd.
Kazuaki Ishizaki, Motohiro Kawahito, Toshiaki Yasue, Mikio Takeuchi, Takeshi Ogasawara, Toshio Suganuma, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
Concurr. Pract. Exp.6
1996 Detection and Global Optimization of Reduction Operations for Distributed Parallel Machines
abstract
This paper presents a new technique for detecting and optimizing reduction operations for parallelizhtg compilers.The technique presented here can detect reduction constructs in general complex loops, parallelize the loops containing reduction constructs, and optimize communications for multiple reduction operations.The optimization proposed here can be applied not only to individual reduction loops, but also to multiple loop nests throughout a program.The techniques have been implemented in our HPF compiler, and their effectiveness is evaluated on an IBM Scalable PowerParallel System SP2 using a set of standard benchmarking programs.Although the aperimental results are still preliminary, it is shown that our techniques for detecting and optimizing reductions are eflective on practical application programs.
Toshio Suganuma, Hideaki Komatsu, Toshio Nakatani
International Conference on Supercomputing1