Tamiya Onodera

dblp:70/1759 · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
0since 2021 · last 2018
0000-0002-6076-8236ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 6 first-authorSystems, architecture and hardware · 6Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
11 papers
Runtime systems and virtual machines · 30% Operating systems · 29% Programming languages and type systems · 15%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 84% Performance modeling and evaluation · 12% Storage systems · 4%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › resource management
memory management
0.222009
Copy-on-write in the PHP language · POPL 2009
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.132006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Stride prefetching by dynamically inspecting objects · PLDI 2003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.132010
A study of Java's non-Java memory · OOPSLA 2010
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
Programming languages and type systems › program equivalence
bisimulation
0.112009
Copy-on-write in the PHP language · POPL 2009
Operating systems › resource management › memory management
copy-on-write
0.112009
Copy-on-write in the PHP language · POPL 2009
Operating systems
interprocess communication
0.112009
Highly scalable web applications with zero-copy data transfer · WWW 2009
Programming languages and type systems
language semantics
0.112009
Copy-on-write in the PHP language · POPL 2009
Operating systems › i/o › i/o subsystem
zero-copy i/o
0.112009
Highly scalable web applications with zero-copy data transfer · WWW 2009
Systems and software security
memory safety
0.112008
Finding bugs in java native interface programs · ISSTA 2008
Runtime systems and virtual machines
garbage collection
0.112008
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Program analysis
static analysis
0.112008
Finding bugs in java native interface programs · ISSTA 2008
Program analysis › type-based analysis
typestate analysis
0.112008
Finding bugs in java native interface programs · ISSTA 2008
Debugging and program repair › software debugging
compiler debugging
0.112006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Concurrent programming
synchronization
0.122002
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
A Study of Locking Objects with Bimodal Fields · OOPSLA 1999
Program analysis
data flow analysis
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Compilers and program optimization › interprocedural optimization
inlining
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Memory systems
cache
0.012003
Stride prefetching by dynamically inspecting objects · PLDI 2003
Memory systems › cache
prefetching
0.012003
Stride prefetching by dynamically inspecting objects · PLDI 2003
Memory systems › cache › prefetching
stride prefetching
0.012003
Stride prefetching by dynamically inspecting objects · PLDI 2003
Concurrent programming › synchronization
locking
0.012002
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
Performance modeling and evaluation
workload characterization
0.012010
A study of Java's non-Java memory · OOPSLA 2010
Internet architecture and protocols › world wide web
web server performance
0.012009
Highly scalable web applications with zero-copy data transfer · WWW 2009
Programming languages and type systems › language implementation
foreign function interface
0.012008
Finding bugs in java native interface programs · ISSTA 2008
Programming languages and type systems › language implementation › foreign function interface
java native interface
0.012008
Finding bugs in java native interface programs · ISSTA 2008
Program analysis › dynamic analysis
profiling
0.012003
Stride prefetching by dynamically inspecting objects · PLDI 2003
Storage systems › data management › database storage
object-oriented database storage
0.011994
Experience with Representing C++ Program Information in an Object-Oriented Database · OOPSLA 1994
Runtime systems and virtual machines › managed runtime
java runtime
0.011999
A Study of Locking Objects with Bimodal Fields · OOPSLA 1999
Program analysis
program representation
0.011994
Experience with Representing C++ Program Information in an Object-Oriented Database · OOPSLA 1994

Methods — techniques the papers use, named apart from their topics

visualization · 0.2memory statistics gathering · 0.2typestate analysis · 0.2syntax checking · 0.2bisimulation · 0.1heap analysis · 0.1garbage collection · 0.1state saving · 0.1replay compilation · 0.1partial interpretation · 0.0object inspection · 0.0exception check elimination · 0.0sharing-oriented clustering · 0.0death-order clustering · 0.0birth-order clustering · 0.0
YearPublicationVenuePosition
2018 Balanced double queues for GC work-stealing on weak memory models
abstract
Work-stealing is promising for scheduling and balancing parallel workloads. It has a wide range of applicability on middleware, libraries, and runtime systems of programming languages. OpenJDK uses work-stealing for copying garbage collection (GC) to balance copying tasks among GC threads. Each thread has its own queue to store tasks. When a thread has no task in its queue, it acts as a thief and attempts to steal a task from another thread's queue. However, this work-stealing algorithm requires expensive memory fences for pushing, popping, and stealing tasks, especially on weak memory models such as POWER and ARM. To address this problem, we propose a work-stealing algorithm that uses double queues. Each GC thread has a public queue that is accessible from other GC threads and a private queue that is only accessible by itself. Pushing and popping tasks in the private queue are free from expensive memory fences. The most significant point in our algorithm is providing a mechanism to maintain the load balance on the basis of the use of double queues. We developed a prototype implementation for parallel GC in OpenJDK8 for ppc64le. We evaluated our algorithm by using SPECjbb2015, SPECjvm2008, TPC-DS, and Apache DayTrader.
Michihiro Horie, Hiroshi Horii, Kazunori Ogata, Tamiya Onodera
ISMM4
2017 Taming Performance Degradation of Containers in the Case of Extreme Memory Overcommitment
abstract
The efficiency of datacenters is important consideration for cloud service providers to make their datacenters always ready for fulfilling the increasing demand for computing resources. Container-based virtualization is one approach to improving efficiency by reducing the overhead of virtualization. Resource overcommitment is another approach, but cloud providers tend to make conservative allocations of resources because there is no good understanding of the relationship between physical resource overcommitment and its impact on performance. This paper presents a quantitative study of performance degradation of containerized workloads due to memory overcommitment and a technique to mitigate it. We focused on physical memory overcommitment, where the sum of the working set memory is larger than the physical memory. We drove a small fraction of Docker containers at a high load level and the rest of them at a very low load level to emulate a common usage pattern of cloud datacenters. Detailed measurements revealed it is difficult to predict how many additional containers can be launched before thrashing hurts performance. We show that tuning the per-container swappiness of heavily loaded containers is effective for launching a larger number of containers and that it achieves an overcommitment of about three times.
Rina Nakazawa, Kazunori Ogata, Seetharami Seelam, Tamiya Onodera
CLOUD4
2016 Workload characterization and optimization of TPC-H queries on Apache Spark
abstract
Besides being an in-memory-oriented computing framework, Spark runs on top of Java Virtual Machines (JVMs), so JVM parameters must be tuned to improve Spark application performance. Misconfigured parameters and settings degrade performance. For example, using Java heaps that are too large often causes a long garbage collection pause time, which accounts for over 10-20% of application execution time. Moreover, recent computing nodes have many cores with simultaneous multi-threading technology and the processors on the node are connected via NUMA, so it is difficult to exploit best performance without taking into account of these hardware features. Thus, optimization in a full stack is also important. Not only JVM parameters but also OS parameters, Spark configuration, and application code based on CPU characteristics need to be optimized to take full advantage of underlying computing resources. In this paper, we used the TPC-H benchmark as our optimization case study and gathered many perspective logs such as application, JVM (e.g. GC and JIT), system utilization, and hardware events from a performance monitoring unit. We discuss current problems and introduce several JVM and OS parameter optimization approaches for accelerating Spark performance. As a result, our optimization exhibits 30-40% increase in speed on average and is up to 5x faster than the naive configuration.
Tatsuhiro Chiba, Tamiya Onodera
ISPASS2
2014 String deduplication for Java-based middleware in virtualized environments
abstract
To increase the memory efficiency in physical servers is a significant concern for increasing the number of virtual machines (VM) in them. When similar web application service runs in each guest VM, many string data with the same values are created in every guest VMs. These duplications of string data are redundant from the viewpoint of memory efficiency in the host OS. This paper proposes two approaches to reduce the duplication in Java string in a single Java VM (JVM) and across JVMs. The first approach is to share string objects cross JVMs by using a read-only memory-mapped file. The other approach is to selectively unify string objects created at runtime in the web applications. This paper evaluates our approach by using the Apache DayTrader and the DaCapo benchmark suite. Our prototype implementation chieved 7% to 12% reduction in the total size of the objects allocated over the lifetime of the programs. In addition, we observed the performance of DayTrader was maintained even under a situation of high density guest VMs in a KVM host machine.
Michihiro Horie, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera
VEE4
2014 A Distributed Quorum System for Ensuring Bounded Staleness of Key-Value Stores
Hiroshi Horii, Miki Enoki, Tamiya Onodera
WAIM3
2013 Increasing the Transparent Page Sharing in Java
abstract
Improving memory utilization is important for improving the efficiency of a cloud datacenter by increasing the number of usable VMs. Memory over-commitment is a common technique for this purpose. Transparent Page Sharing (TPS) is a technique to improve the utilization by sharing identical memory pages to reduce the total memory consumption. For a cloud datacenter, we might expect TPS will reduce memory usage because VMs often execute the same OS and middleware and thus they may have many identical pages. However, TPS is less effective for Java-based middleware because the Java VM finds it difficult to manage the layouts of internal data structures that depend on the execution of Java programs. This paper presents detailed breakdowns of the memory usage of KVM guest VMs executing a Java-based Web application server. Then we propose increasing the amount of page sharing by utilizing a class sharing mechanism in the Java VM. Our approach reduced the measured physical memory for class metadata by up to 89.6% when using the Apache DayTrader benchmark running on four guest VMs in a KVM host machine.
Kazunori Ogata, Tamiya Onodera
ISPASS2
2012 Memory-Efficient Index for Cache Invalidation Mechanism with OpenJPA
Miki Enoki, Yosuke Ozawa, Hiroshi Horii, Tamiya Onodera
WISE4
2010 Performance Improvement of OpenJPA by Query Dependency Analysis
Miki Enoki, Yosuke Ozawa, Tamiya Onodera
DASFAA (2)3
2010 A study of Java's non-Java memory
abstract
A Java application sometimes raises an out-of-memory ex-ception. This is usually because it has exhausted the Java heap. However, a Java application can raise an out-of-memory exception when it exhausts the memory used by Java that is not in the Java heap. We call this area non-Java memory. For example, an out-of-memory exception in the non-Java memory can happen when the JVM attempts to load too many classes. Although it is relatively rare to ex-haust the non-Java memory compared to exhausting the Java heap, a Java application can consume a considerable amount of non-Java memory.This paper presents a quantitative analysis of non-Java memory. To the best of our knowledge, this is the first in-depth analysis of the non-Java memory. To do this we cre-ated a tool called Memory Analyzer for Redundant, Unused, and String Areas (MARUSA), which gathers memory statis-tics from both the OS and the Java virtual machine, break-ing down and visualizing the non-Java memory usage.We studied the use of non-Java memory for a wide range of Java applications, including the DaCapo benchmarks and Apache DayTrader. Our study is based on the IBM J9 Java Virtual Machine for Linux. Although some of our results may be specific to this combination, we believe that most of our observations are applicable to other platforms as well.
Kazunori Ogata, Dai Mikurube, Kiyokuni Kawachiya, Scott Trent, Tamiya Onodera
OOPSLA5
2010 Scalable performance of system S for extract-transform-load processing
abstract
ETL (Extract-Transform-Load) processing is filling an increasingly critical role in analyzing business data and in taking appropriate business actions based on the results. As the volume of business data to be analyzed increases and quick responses are more critical for business success, there are strong demands for scalable high-performance ETL processors. In this paper, we evaluate a distributed data stream processing engine called System S for those purposes. Based on the original motivation of building System S as a data stream processing engine, we first perform a qualitative study to see if the programming model of System S is suitable for representing an ETL workflow. Second we did performance studies with a representative ETL scenario. Through our series of experiments, we found that the SPADE programming model and its runtime environment naturally fits the requirements of handling massive amounts of ETL data in a highly scalable manner.
Toyotaro Suzumura, Toshiaki Yasue, Tamiya Onodera
SYSTOR3
2010 Efficient runtime tracking of allocation sites in Java
abstract
Tracking the allocation site of every object at runtime is useful for reliable, optimized Java. To be used in production environments, the tracking must be accurate with minimal speed loss. Previous approaches suffer from performance degradation due to the additional field added to each object or track the allocation sites only probabilistically. We propose two novel approaches to track the allocation sites of every object in Java with only a 1.0% slow-down on average. Our first approach, the Allocation-Site-as-a-Hash-code (ASH) Tracker, encodes the allocation site ID of an object into the hash code field of its header by regarding the ID as part of the hash code. ASH Tracker avoids an excessive increase in hash code collisions by dynamically shrinking the bit-length of the ID as more and more objects are allocated at that site. For those Java VMs without the hash code field, our second approach, the Allocation-Site-via-a-Class-pointer (ASC) Tracker, makes the class pointer field in an object header refer to the allocation site structure of the object, which in turn points to the actual class structure. ASC Tracker mitigates the indirection overhead by constant-class-field duplication and allocation-site equality checks. While a previous approach of adding a 4-byte field caused up to 14.4% and an average 5% slowdown, both ASH and ASC Trackers incur at most a 2.0% and an average 1.0% loss. We demonstrate the usefulness of our low-overhead trackers by an allocation-site-aware memory leak detector and allocation-site-based pretenuring in generational GC. Our pretenuring achieved on average 1.8% and up to 11.8% speedups in SPECjvm2008.
Rei Odaira, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera, Toshio Nakatani
VEE4
2010 Evaluation of a just-in-time compiler retrofitted for PHP
abstract
Programmers who develop Web applications often use dynamic scripting languages such as Perl, PHP, Python, and Ruby. For general purpose scripting language usage, interpreter-based implementations are efficient and popular but the server-side usage for Web application development implies an opportunity to significantly enhance Web server throughput. This paper summarizes a study of the optimization of PHP script processing. We developed a PHP processor, P9, by adapting an existing production-quality just-in-time (JIT) compiler for a Java virtual machine, for which optimization technologies have been well-established, especially for server-side application. This paper describes and contrasts microbenchmarks and SPECweb2005 benchmark results for a well-tuned configuration of a traditional PHP interpreter and our JIT compiler-based implementation, P9. Experimental results with the microbenchmarks show 2.5-9.5x advantage with P9, and the SPECweb2005 measurements show 20-30 % improvements. These results show that the acceleration of dynamic scripting language processing does matter in a realistic Web application server environment. CPU usage profiling shows our simple JIT compiler introduction reduces the PHP core runtime overhead from 45 % to 13 % for a SPECweb2005 scenario, implying that further improvements of dynamic compilers would provide little additional return unless other major overheads such as heavy memory copy between the language runtime and Web server frontend are reduced.
Michiaki Tatsubori, Akihiko Tozawa, Toyotaro Suzumura, Scott Trent, Tamiya Onodera
VEE5
2009 Copy-on-write in the PHP language
abstract
PHP is a popular language for server-side applications. In PHP, assignment to variables copies the assigned values, according to its so-called copy-on-assignment semantics. In contrast, a typical PHP implementation uses a copy-on-write scheme to reduce the copy overhead by delaying copies as much as possible. This leads us to ask if the semantics and implementation of PHP coincide, and actually this is not the case in the presence of sharings within values. In this paper, we describe the copy-on-assignment semantics with three possible strategies to copy values containing sharings. The current PHP implementation has inconsistencies with these semantics, caused by its naïve use of copy-on-write. We fix this problem by the novel mostly copy-on-write scheme, making the copy-on-write implementations faithful to the semantics. We prove that our copy-on-write implementations are correct, using bisimulation with the copy-on-assignment semantics.
Akihiko Tozawa, Michiaki Tatsubori, Tamiya Onodera, Yasuhiko Minamide
POPL3
2009 Highly scalable web applications with zero-copy data transfer
abstract
The performance of server-side applications is becoming increasingly important as more applications exploit the Web application model. Extensive work has been done to improve the performance of individual software components such as Web servers and programming language runtimes. This paper describes a novel approach to boost Web application performance by improving inter-process communication between a programming language runtime and Web server runtime. The approach reduces redundant processing for memory copying and the context switch overhead between user space and kernel space by exploiting the zero-copy data transfer methodology, such as the sendfile system call. In order to transparently utilize this optimization feature with existing Web applications, we propose enhancements of the PHP runtime, FastCGI protocol, and Web server. Our proposed approach achieves a 126% performance improvement with micro-benchmarks and a 44% performance improvement for a standard Web benchmark, SPECweb2005.
Toyotaro Suzumura, Michiaki Tatsubori, Scott Trent, Akihiko Tozawa, Tamiya Onodera
WWW5
2008 Performance Comparison of Web Service Engines in PHP, Java and C
abstract
PHP is well known as a programming language in the Web 2.0 era enabling agile server-side software development. It has officially supported SOAP messaging since version 5 through a C-based built-in library. In this paper we perform a thorough study of the capability of PHP as a Web service engine in both qualitative and quantitative aspects while comparing it with other Web service engines implemented in Java and C. We used Axis2 for this purpose as it is an open source web service engine whose implementation is available both in Java and C. We report that PHP as a web service engine performs competitively with Axis2 Java for Web services involving small payloads, and greatly outperforms it for larger payloads by 5-17 times. As the authors expected, Axis2 C performs best, but the experimental results demonstrate that PHP performance is closer to Axis2 C with larger payloads. This performance difference comes from the fact that the SOAP engine within the PHP runtime is implemented in C with a monolithic architecture, whereas Axis2 uses a more modular architecture for the flexible insertation of handlers for an assorted set of WS-* standards, and also that Axis2 uses a different data binding mechanism known as ADB (Axis2 Data binding). This paper is the first attempt to compare Web services engines implemented in PHP, Java and C, and the authors believe that this boosts the development of SOAP-based Web services in PHP by letting people know its decent performance score and high productivity characteristics.
Toyotaro Suzumura, Scott Trent, Michiaki Tatsubori, Akihiko Tozawa, Tamiya Onodera
ICWS5
2008 Finding bugs in java native interface programs
abstract
In this paper, we describe static analysis techniques for finding bugs in programs using the Java Native Interface (JNI). The JNI is both tedious and error-prone because there are many JNI-specific mistakes that are not caught by a native compiler. This paper is focused on four kinds of common mistakes. First, explicit statements to handle a possible exception need to be inserted after a statement calling a Java method. However, such statements tend to be forgotten. We present a typestate analysis to detect this exception-handling mistake. Second, while the native code can allocate resources in a Java VM, those resources must be manually released, unlike Java. Mistakes in resource management cause leaks and other errors. To detect Java resource errors, we used the typestate analysis also used for detecting general memory errors. Third, if a reference to a Java resource lives across multiple native method invocations, it should be converted into a global reference. However, programmers sometimes forget this rule and, for example, store a local reference in a global variable for later uses. We provide a syntax checker that detects this bad coding practice. Fourth, no JNI function should be called in a critical region. If called there, the current thread might block and cause a deadlock. Misinterpreting the end of the critical region, programmers occasionally break this rule. We present a simple typestate analysis to detect an improper JNI function call in a critical region.
Goh Kondoh, Tamiya Onodera
ISSTA2
2008 Performance Comparison of PHP and JSP as Server-Side Scripting Languages
Scott Trent, Michiaki Tatsubori, Toyotaro Suzumura, Akihiko Tozawa, Tamiya Onodera
Middleware5
2008 Analysis and reduction of memory inefficiencies in Java strings
abstract
This paper describes a novel approach to reduce the memory consumption of Java programs, by focusing on their "string memory inefficiencies". In recent Java applications, string data occupies a large amount of the heap area. For example, about 40% of the live heap area is used for string data when a production J2EE application server is running. By investigating the string data in the live heap, we identified two types of memory inefficiencies -- "duplication" and "unused literals". In the heap, there are many string objects that have the same values. There also exist many string literals whose values are not actually used by the application. Since these inefficiencies exist as live objects, they cannot be eliminated by existing garbage collection techniques, which only remove dead objects. Quantitative analysis of Java heaps in real applications revealed that more than 50% of the string data in the live heap is wasted by these inefficiencies. To reduce the string memory inefficiencies, this paper proposes two techniques at the Java virtual machine level, "StringGC" for eliminating duplicated strings at the time of garbage collection, and "Lazy Body Creation" for delaying part of the literal instantiation until the literal's value is actually used. We also present an interesting technique at the Java program level, which we call "BundleConverter", for preventing unused message literals from being instantiated. Prototype implementations on a production Java virtual machine have achieved about 18% reduction of the live heap in the production application server. The proposed techniques could also reduce the live heap of standard Java benchmarks by 11.6% on average, without noticeable performance degradation.
Kiyokuni Kawachiya, Kazunori Ogata, Tamiya Onodera
OOPSLA3
2007 Cloneable JVM: a new approach to start isolated java applications faster
abstract
Java has been successful particularly for writing applications in the server environment. However, isolation of multiple applications hasnot been efficiently achieved in Java. Many customers require that their applications are guarded by independent OS processes, but starting a Java application with a new process results in a long sequence of initializations being repeated each time. To date, there has been no way to quickly start a new Java application as an isolated OS process. In this paper, we propose a new isolation approach called Cloneable JVM to eliminate this startup overhead in Java. The key idea is to createa new Java application by copying, or cloning, the already-initialized image of the primary JVM process. Since the clone is already initialized, it can begin actual operations immediately as a new isolated process. This cloning abstraction can support new scenarios for Java, such as user isolation and transaction isolation. We implemented a prototype of the Cloneable JVM by modifying a production JVM on Linux, which provides a new API for cloning constructed on the Isolate API defined in JSR 121. Using this cloning API, several Java applications, including a large production J2EE application server, we remodified to demonstrate the isolation scenarios. Evaluations using these prototypes showed that new ready-to-serve Java applications can start up as a new process in less than 5 seconds, which is 4 to 170 times faster than starting these applications from scratch.
Kiyokuni Kawachiya, Kazunori Ogata, Daniel Silva 0001, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
VEE4
2006 Replay compilation: improving debuggability of a just-in-time compiler
abstract
The performance of Java has been tremendously improved by the advance of Just-in-Time (JIT) compilation technologies. However, debugging such a dynamic compiler is much harder than a static compiler. Recompiling the problematic method to produce a diagnostic output does not necessarily work as expected, because the compilation of a method depends on runtime information at the time of compilation.In this paper, we propose a new approach, called replay JIT compilation, which can reproduce the same compilation remotely by using two compilers, the state-saving compiler and the replaying compiler. The state-saving compiler is used in a normal run, and, while compiling a method, records into a log all of the input for the compiler. The replaying compiler is then used in a debugging run with the system dump, to recompile a method with the options for diagnostic output. We reduced the overhead to save the input by using the system dump and by categorizing the input based on how its value changes. In our experiment, the increase of the compilation time for saving the input was only 1%, and the size of the additional memory needed for saving the input was only 10% of the compiler-generated code.
Kazunori Ogata, Tamiya Onodera, Kiyokuni Kawachiya, Hideaki Komatsu, Toshio Nakatani
OOPSLA2
2004 Lock Reservation for Java Reconsidered
Tamiya Onodera, Kiyokuni Kawachiya, Akira Koseki
ECOOP1
2003 Effectiveness of cross-platform optimizations for a java just-in-time compiler
abstract
This paper describes the system overview of our Java Just-In-Time (JIT) compiler, which is the basis for the latest production version of IBM Java JIT compiler that supports a diversity of processor architectures including both 32-bit and 64-bit modes, CISC, RISC, and VLIW architectures. In particular, we focus on the design and evaluation of the cross-platform optimizations that are common across different architectures. We studied the effectiveness of each optimization by selectively disabling it in our JIT compiler on three different platforms: IA-32, IA-64, and PowerPC. Our detailed measurements allowed us to rank the optimizations in terms of the greatest performance improvements with the smallest compilation times. The identified set includes method inlining only for tiny methods, exception check eliminations using forward dataflow analysis and partial redundancy elimination, scalar replacement for instance and class fields using dataflow analysis, optimizations for type inclusion checks, and the elimination of merge points in the control flow graphs. These optimizations can achieve 90% of the peak performance for two industry-standard benchmark programs on these platforms with only 34% of the compilation time compared to the case for using all of the optimizations.
Kazuaki Ishizaki, Mikio Takeuchi, Kiyokuni Kawachiya, Toshio Suganuma, Osamu Gohda, Tatsushi Inagaki, Akira Koseki, Kazunori Ogata, Motohiro Kawahito, Toshiaki Yasue, Takeshi Ogasawara, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
OOPSLA12
2003 Stride prefetching by dynamically inspecting objects
abstract
Software prefetching is a promising technique to hide cache miss latencies, but it remains challenging to effectively prefetch pointer-based data structures because obtaining the memory address to be prefetched requires pointer dereferences. The recently proposed stride prefetching overcomes this problem, but it only exploits inter-iteration stride patterns and relies on an off-line profiling method.We propose a new algorithm for stride prefetching which is intended for use in a dynamic compiler. We exploit both inter- and intra-iteration stride patterns, which we discover using an ultra-lightweight profiling technique, called object inspection. This is a kind of partial interpretation that only a dynamic compiler can perform. During the compilation of a method, the dynamic compiler gathers the profile information by partially interpreting the method using the actual values of parameters and causing no side effects.We evaluated an implementation of our prefetching algorithm in a production-level Java just-in time compiler. The results show that the algorithm achieved up to an 18.9% and 25.1% speedup in industry-standard benchmarks on the Pentium 4 and the Athlon MP, respectively, while it increased the compilation time by less than 3.0%.
Tatsushi Inagaki, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
PLDI2
2002 Lock reservation: Java locks can mostly do without atomic operations
abstract
Because of the built-in support for multi-threaded programming, Java programs perform many lock operations. Although the overhead has been significantly reduced in the recent virtual machines, One or more atomic operations are required for acquiring and releasing an object's lock even in the fastest cases.This paper presents a novel algorithm called lock reservation. It exploits thread locality of Java locks, which claims that the locking sequence of a Java lock contains a very long repetition of a specific thread. The algorithm allows locks to be reserved for threads. When a thread attempts to acquire a lock, it can do without any atomic operation if the lock is reserved for the thread. Otherwise, it cancels the reservation and falls back to a conventional locking algorithm.We have evaluated an implementation of lock reservation in IBM's production virtual machine and compiler. The results show that it achieved performance improvements up to 53% in real Java programs.
Kiyokuni Kawachiya, Akira Koseki, Tamiya Onodera
OOPSLA3
2000 Design, implementation, and evaluation of optimizations in a JavaTM Just-In-Time compiler
abstract
The Java language incurs a runtime overhead for exception checks and object accesses, which are executed without an interior pointer in order to ensure safety. It also requires type inclusion test, dynamic class loading, and dynamic method calls in order to ensure flexibility. A ‘Just-In-Time’ (JIT) compiler generates native code from Java byte code at runtime. It must improve the runtime performance without compromising the safety and flexibility of the Java language. We designed and implemented effective optimizations for a JIT compiler, such as exception check elimination, common subexpression elimination, simple type inclusion test, method inlining, and devirtualization of dynamic method call. We evaluate the performance benefits of these optimizations based on various statistics collected using SPECjvm98, its candidates, and two JavaSoft applications with byte code sizes ranging from 23 000 to 280 000 bytes. Each optimization contributes to an improvement in the performance of the programs. Copyright © 2000 John Wiley & Sons, Ltd.
Kazuaki Ishizaki, Motohiro Kawahito, Toshiaki Yasue, Mikio Takeuchi, Takeshi Ogasawara, Toshio Suganuma, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
Concurr. Pract. Exp.7
1999 A Study of Locking Objects with Bimodal Fields
abstract
Object locking can be efficiently implemented by bimodal use of a field reserved in an object. The field is used as a lightweight lock in one mode, while it holds a reference to a heavyweight lock in the other mode. A bimodal locking algorithm recently proposed for Java achieves the highest performance in the absence of contention, and is still fast enough when contention occurs.
Tamiya Onodera, Kiyokuni Kawachiya
OOPSLA1
1997 Optimizing Smalltalk by Selector Code INdexing Can Be Practical
Tamiya Onodera, Hiroaki Nakamura
ECOOP1
1994 Experience with Representing C++ Program Information in an Object-Oriented Database
abstract
Two major issues related to storing program information in an OODB are sharing and clustering. The former is important since it prevents the database from consuming excessive disk space, while the latter is crucial, since it keeps clients running without thrashing. In our database, objects are shared across multiple programs' translation units, and are clustered by combining three techniques, namely, birth-order, death-order, and sharing-oriented clusterings. An initial experiment shows that, for a medium-size application, the database consumes 3.5 times less disk space than in a conventional environment, and that the invocation of a client is almost instantaneous.
Tamiya Onodera
OOPSLA1
1993 Reducing Compilation Time by a Compilation Server
abstract
Abstract In language systems that support separate compilation, we often observe that header files are internalized over and over again when the source files that depend on them are compiled. Making a compiler a long‐lived server eliminates such redundant processing of header files, thus reducing the compilation time. The paper first describes compilation servers for C‐family languages in general, and then a compilation server for our C‐based object‐oriented language in particular. The performance results of our server show that a compilation server can substantially shorten the compilation time.
Tamiya Onodera
Softw. Pract. Exp.1
1993 A Generational and Conservative Copying Collector for Hybrid Objectoriented Languages
abstract
Abstract A copying collector has two excellent properties: it compacts the heap, and the execution time depends solely on the number of live objects. Use of a copying collector is thought by some to be a more efficient way of managing the heap than explicit freeing of objects. This paper describes a high‐performance copying collector for a hybrid object‐oriented language. The collector is both conservative and generational. It relies on the overlying compiler to identify most true pointers, and on the underlying operating system to detect pointers to younger generations. The implementation described here uses a modified version of the compiler for a C‐based object‐oriented language, and the Mach operating system. The performance results have confirmed the author's expectation: the collector has been faster than explicit freeing.
Tamiya Onodera
Softw. Pract. Exp.1
1986 A formalization for the specification and systematic generation of computer graphics systems
Tamiya Onodera, Satoru Kawai
Vis. Comput.1