Kazunori Ogata

dblp:45/2673 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 4 first-authorSystems, architecture and hardware · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Runtime systems and virtual machines · 63% Operating systems · 13% Debugging and program repair · 10%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 72% Performance modeling and evaluation · 21% Processor architecture and microarchitecture · 7%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.122010
A study of Java's non-Java memory · OOPSLA 2010
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.122006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Runtime systems and virtual machines
garbage collection
0.112008
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Operating systems › resource management
memory management
0.112008
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Debugging and program repair › software debugging
compiler debugging
0.112006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Program analysis
data flow analysis
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Compilers and program optimization › interprocedural optimization
inlining
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Runtime systems and virtual machines › interpreter
interpreter optimization
0.012002
Bytecode fetch optimization for a Java interpreter · ASPLOS 2002
Runtime systems and virtual machines › interpreter
interpreter performance
0.012002
Bytecode fetch optimization for a Java interpreter · ASPLOS 2002
Performance modeling and evaluation
workload characterization
0.012010
A study of Java's non-Java memory · OOPSLA 2010
Processor architecture and microarchitecture
superscalar processor
0.012002
Bytecode fetch optimization for a Java interpreter · ASPLOS 2002

Methods — techniques the papers use, named apart from their topics

visualization · 0.2memory statistics gathering · 0.2heap analysis · 0.1garbage collection · 0.1speculative decoding · 0.1handler customization · 0.1state saving · 0.1replay compilation · 0.1partial redundancy elimination · 0.0exception check elimination · 0.0top-of-stack caching · 0.0
YearPublicationVenuePosition
2019 Gcom-C/Sgli Ocean Standard Products and Early Validation Results
abstract
GCOM-C/SGLI is a multi-wavelength optical radiometer launched on December 23, 2017. The data provision has started from December 20, 2018. In this research, we briefly introduce standard Level 2 products and their validation results based on in situ data. Mean absolute percentage differences are 16.3 - 69.6% for NLW between 380 - 670 nm, 39 and 64.5% for AOT at 670 nm and 865 nm, 27.9% for CHL and 44.2% for aCDOM. Although the number of in situ validation data are still scare for a few ocean color products, the accuracy of GCOM-C/SGLI data will be improved by our future efforts of calibration/validation activities.
Mitsuhiro Toratani, Stanford B. Hooker, Yoko Kiyomoto, Hiroshi Murakami, Yukio Kurihara, Masahiro Hori, Hisatomo Waga, Youhei Yamashita, Akihiko Tanaka, Kazunori Ogata, Koji Suzuki, Joji Ishizaka, Toru Hirawake, Takafumi Hirata, Tomonori Isada, Hiroto Higa, Victor S. Kuwahara
IGARSS10
2019 Scaling up parallel GC work-stealing in many-core environments
abstract
Parallel copying garbage collection (GC) is widely used in the de facto Java virtual machines such as OpenJDK and OpenJ9. OpenJDK uses work-stealing for copying objects in the Parallel GC and Garbage-First (G1) GC policies to balance the copying task among GC threads. When a thread has no task in its own queue, it tries to steal a task from another thread's queue as a thief. When a thief succeeds in stealing a task, it processes the task and enqueues the children of the task into its queue, which is accessible from other thieves.Unfortunately, the overhead of the work-stealing framework becomes non-negligible when we aim to achieve a minimum GC pause time by increasing the number of GC threads. Since the number of tasks processed per thread decreases, thieves frequently try to steal tasks from others at a low success rate. When a thief fails in steals continuously, it needs to wait in a spin loop on the termination protocol of the work-stealing framework. Spinning in a loop frequently results in high CPU utilization, which is not acceptable in a large-scale data center where severe power management is required. This paper proposes two approaches named steal-best-of-many selection and spin-less termination to reduce the overhead in the work-stealing framework. Steal-best-of-many selection reduces steal failures by changing the number of queue selections to steal in accordance with the number of GC threads. Spin-less termination moves a part of the object copies into a spin loop by changing the procedure of copying GC. It reduces part of the GC pause time for the object copy as well as the CPU utilization for the spin loop. We developed a prototype on OpenJDK8 and evaluated it using SPECjbb2015 and SPECjvm2008 benchmarks. Critical-jOPS performance of SPECjbb2015 improved by 18% at maximum and scores of the SPECjvm2008 benchmarks improved by 1-5%.
Michihiro Horie, Kazunori Ogata, Mikio Takeuchi, Hiroshi Horii
ISMM2
2018 Balanced double queues for GC work-stealing on weak memory models
abstract
Work-stealing is promising for scheduling and balancing parallel workloads. It has a wide range of applicability on middleware, libraries, and runtime systems of programming languages. OpenJDK uses work-stealing for copying garbage collection (GC) to balance copying tasks among GC threads. Each thread has its own queue to store tasks. When a thread has no task in its queue, it acts as a thief and attempts to steal a task from another thread's queue. However, this work-stealing algorithm requires expensive memory fences for pushing, popping, and stealing tasks, especially on weak memory models such as POWER and ARM. To address this problem, we propose a work-stealing algorithm that uses double queues. Each GC thread has a public queue that is accessible from other GC threads and a private queue that is only accessible by itself. Pushing and popping tasks in the private queue are free from expensive memory fences. The most significant point in our algorithm is providing a mechanism to maintain the load balance on the basis of the use of double queues. We developed a prototype implementation for parallel GC in OpenJDK8 for ppc64le. We evaluated our algorithm by using SPECjbb2015, SPECjvm2008, TPC-DS, and Apache DayTrader.
Michihiro Horie, Hiroshi Horii, Kazunori Ogata, Tamiya Onodera
ISMM3
2017 Taming Performance Degradation of Containers in the Case of Extreme Memory Overcommitment
abstract
The efficiency of datacenters is important consideration for cloud service providers to make their datacenters always ready for fulfilling the increasing demand for computing resources. Container-based virtualization is one approach to improving efficiency by reducing the overhead of virtualization. Resource overcommitment is another approach, but cloud providers tend to make conservative allocations of resources because there is no good understanding of the relationship between physical resource overcommitment and its impact on performance. This paper presents a quantitative study of performance degradation of containerized workloads due to memory overcommitment and a technique to mitigate it. We focused on physical memory overcommitment, where the sum of the working set memory is larger than the physical memory. We drove a small fraction of Docker containers at a high load level and the rest of them at a very low load level to emulate a common usage pattern of cloud datacenters. Detailed measurements revealed it is difficult to predict how many additional containers can be launched before thrashing hurts performance. We show that tuning the per-container swappiness of heavily loaded containers is effective for launching a larger number of containers and that it achieves an overcommitment of about three times.
Rina Nakazawa, Kazunori Ogata, Seetharami Seelam, Tamiya Onodera
CLOUD2
2017 GCOM-C/SGLI Level-2 ocean color products generation
abstract
The SGLI instrument, which is to be launched by JAXA in March 2017 aboard GCOM-C satellite, conducts ocean color observation in 250 m spatial resolution (SST in 500 m resolution) over coastal region, in addition to the global ocean observation in 1 km resolution. We first describe here the SGLI Level 2 (L2) ocean standard products, which consist of 6 ocean color parameters and 1 SST parameter, and then briefly describe the product generation data flow together with the algorithms implemented in the production system. The required time and memory size of the Level-2 data generation was evaluated over a simulated SGLI data scene, which indicates that the modules meet the standard L2 data generation system requirements, although further optimization in the processing code will be pursued.
Kazunori Ogata, Mitsuhiro Toratani, Hiroshi Murakami
IGARSS1
2014 String deduplication for Java-based middleware in virtualized environments
abstract
To increase the memory efficiency in physical servers is a significant concern for increasing the number of virtual machines (VM) in them. When similar web application service runs in each guest VM, many string data with the same values are created in every guest VMs. These duplications of string data are redundant from the viewpoint of memory efficiency in the host OS. This paper proposes two approaches to reduce the duplication in Java string in a single Java VM (JVM) and across JVMs. The first approach is to share string objects cross JVMs by using a read-only memory-mapped file. The other approach is to selectively unify string objects created at runtime in the web applications. This paper evaluates our approach by using the Apache DayTrader and the DaCapo benchmark suite. Our prototype implementation chieved 7% to 12% reduction in the total size of the objects allocated over the lifetime of the programs. In addition, we observed the performance of DayTrader was maintained even under a situation of high density guest VMs in a KVM host machine.
Michihiro Horie, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera
VEE2
2013 Increasing the Transparent Page Sharing in Java
abstract
Improving memory utilization is important for improving the efficiency of a cloud datacenter by increasing the number of usable VMs. Memory over-commitment is a common technique for this purpose. Transparent Page Sharing (TPS) is a technique to improve the utilization by sharing identical memory pages to reduce the total memory consumption. For a cloud datacenter, we might expect TPS will reduce memory usage because VMs often execute the same OS and middleware and thus they may have many identical pages. However, TPS is less effective for Java-based middleware because the Java VM finds it difficult to manage the layouts of internal data structures that depend on the execution of Java programs. This paper presents detailed breakdowns of the memory usage of KVM guest VMs executing a Java-based Web application server. Then we propose increasing the amount of page sharing by utilizing a class sharing mechanism in the Java VM. Our approach reduced the measured physical memory for class metadata by up to 89.6% when using the Apache DayTrader benchmark running on four guest VMs in a KVM host machine.
Kazunori Ogata, Tamiya Onodera
ISPASS1
2010 A study of Java's non-Java memory
abstract
A Java application sometimes raises an out-of-memory ex-ception. This is usually because it has exhausted the Java heap. However, a Java application can raise an out-of-memory exception when it exhausts the memory used by Java that is not in the Java heap. We call this area non-Java memory. For example, an out-of-memory exception in the non-Java memory can happen when the JVM attempts to load too many classes. Although it is relatively rare to ex-haust the non-Java memory compared to exhausting the Java heap, a Java application can consume a considerable amount of non-Java memory.This paper presents a quantitative analysis of non-Java memory. To the best of our knowledge, this is the first in-depth analysis of the non-Java memory. To do this we cre-ated a tool called Memory Analyzer for Redundant, Unused, and String Areas (MARUSA), which gathers memory statis-tics from both the OS and the Java virtual machine, break-ing down and visualizing the non-Java memory usage.We studied the use of non-Java memory for a wide range of Java applications, including the DaCapo benchmarks and Apache DayTrader. Our study is based on the IBM J9 Java Virtual Machine for Linux. Although some of our results may be specific to this combination, we believe that most of our observations are applicable to other platforms as well.
Kazunori Ogata, Dai Mikurube, Kiyokuni Kawachiya, Scott Trent, Tamiya Onodera
OOPSLA1
2010 Efficient runtime tracking of allocation sites in Java
abstract
Tracking the allocation site of every object at runtime is useful for reliable, optimized Java. To be used in production environments, the tracking must be accurate with minimal speed loss. Previous approaches suffer from performance degradation due to the additional field added to each object or track the allocation sites only probabilistically. We propose two novel approaches to track the allocation sites of every object in Java with only a 1.0% slow-down on average. Our first approach, the Allocation-Site-as-a-Hash-code (ASH) Tracker, encodes the allocation site ID of an object into the hash code field of its header by regarding the ID as part of the hash code. ASH Tracker avoids an excessive increase in hash code collisions by dynamically shrinking the bit-length of the ID as more and more objects are allocated at that site. For those Java VMs without the hash code field, our second approach, the Allocation-Site-via-a-Class-pointer (ASC) Tracker, makes the class pointer field in an object header refer to the allocation site structure of the object, which in turn points to the actual class structure. ASC Tracker mitigates the indirection overhead by constant-class-field duplication and allocation-site equality checks. While a previous approach of adding a 4-byte field caused up to 14.4% and an average 5% slowdown, both ASH and ASC Trackers incur at most a 2.0% and an average 1.0% loss. We demonstrate the usefulness of our low-overhead trackers by an allocation-site-aware memory leak detector and allocation-site-based pretenuring in generational GC. Our pretenuring achieved on average 1.8% and up to 11.8% speedups in SPECjvm2008.
Rei Odaira, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera, Toshio Nakatani
VEE2
2008 Analysis and reduction of memory inefficiencies in Java strings
abstract
This paper describes a novel approach to reduce the memory consumption of Java programs, by focusing on their "string memory inefficiencies". In recent Java applications, string data occupies a large amount of the heap area. For example, about 40% of the live heap area is used for string data when a production J2EE application server is running. By investigating the string data in the live heap, we identified two types of memory inefficiencies -- "duplication" and "unused literals". In the heap, there are many string objects that have the same values. There also exist many string literals whose values are not actually used by the application. Since these inefficiencies exist as live objects, they cannot be eliminated by existing garbage collection techniques, which only remove dead objects. Quantitative analysis of Java heaps in real applications revealed that more than 50% of the string data in the live heap is wasted by these inefficiencies. To reduce the string memory inefficiencies, this paper proposes two techniques at the Java virtual machine level, "StringGC" for eliminating duplicated strings at the time of garbage collection, and "Lazy Body Creation" for delaying part of the literal instantiation until the literal's value is actually used. We also present an interesting technique at the Java program level, which we call "BundleConverter", for preventing unused message literals from being instantiated. Prototype implementations on a production Java virtual machine have achieved about 18% reduction of the live heap in the production application server. The proposed techniques could also reduce the live heap of standard Java benchmarks by 11.6% on average, without noticeable performance degradation.
Kiyokuni Kawachiya, Kazunori Ogata, Tamiya Onodera
OOPSLA2
2007 Cloneable JVM: a new approach to start isolated java applications faster
abstract
Java has been successful particularly for writing applications in the server environment. However, isolation of multiple applications hasnot been efficiently achieved in Java. Many customers require that their applications are guarded by independent OS processes, but starting a Java application with a new process results in a long sequence of initializations being repeated each time. To date, there has been no way to quickly start a new Java application as an isolated OS process. In this paper, we propose a new isolation approach called Cloneable JVM to eliminate this startup overhead in Java. The key idea is to createa new Java application by copying, or cloning, the already-initialized image of the primary JVM process. Since the clone is already initialized, it can begin actual operations immediately as a new isolated process. This cloning abstraction can support new scenarios for Java, such as user isolation and transaction isolation. We implemented a prototype of the Cloneable JVM by modifying a production JVM on Linux, which provides a new API for cloning constructed on the Isolate API defined in JSR 121. Using this cloning API, several Java applications, including a large production J2EE application server, we remodified to demonstrate the isolation scenarios. Evaluations using these prototypes showed that new ready-to-serve Java applications can start up as a new process in less than 5 seconds, which is 4 to 170 times faster than starting these applications from scratch.
Kiyokuni Kawachiya, Kazunori Ogata, Daniel Silva 0001, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
VEE2
2006 Replay compilation: improving debuggability of a just-in-time compiler
abstract
The performance of Java has been tremendously improved by the advance of Just-in-Time (JIT) compilation technologies. However, debugging such a dynamic compiler is much harder than a static compiler. Recompiling the problematic method to produce a diagnostic output does not necessarily work as expected, because the compilation of a method depends on runtime information at the time of compilation.In this paper, we propose a new approach, called replay JIT compilation, which can reproduce the same compilation remotely by using two compilers, the state-saving compiler and the replaying compiler. The state-saving compiler is used in a normal run, and, while compiling a method, records into a log all of the input for the compiler. The replaying compiler is then used in a debugging run with the system dump, to recompile a method with the options for diagnostic output. We reduced the overhead to save the input by using the system dump and by categorizing the input based on how its value changes. In our experiment, the increase of the compilation time for saving the input was only 1%, and the size of the additional memory needed for saving the input was only 10% of the compiler-generated code.
Kazunori Ogata, Tamiya Onodera, Kiyokuni Kawachiya, Hideaki Komatsu, Toshio Nakatani
OOPSLA1
2003 Effectiveness of cross-platform optimizations for a java just-in-time compiler
abstract
This paper describes the system overview of our Java Just-In-Time (JIT) compiler, which is the basis for the latest production version of IBM Java JIT compiler that supports a diversity of processor architectures including both 32-bit and 64-bit modes, CISC, RISC, and VLIW architectures. In particular, we focus on the design and evaluation of the cross-platform optimizations that are common across different architectures. We studied the effectiveness of each optimization by selectively disabling it in our JIT compiler on three different platforms: IA-32, IA-64, and PowerPC. Our detailed measurements allowed us to rank the optimizations in terms of the greatest performance improvements with the smallest compilation times. The identified set includes method inlining only for tiny methods, exception check eliminations using forward dataflow analysis and partial redundancy elimination, scalar replacement for instance and class fields using dataflow analysis, optimizations for type inclusion checks, and the elimination of merge points in the control flow graphs. These optimizations can achieve 90% of the peak performance for two industry-standard benchmark programs on these platforms with only 34% of the compilation time compared to the case for using all of the optimizations.
Kazuaki Ishizaki, Mikio Takeuchi, Kiyokuni Kawachiya, Toshio Suganuma, Osamu Gohda, Tatsushi Inagaki, Akira Koseki, Kazunori Ogata, Motohiro Kawahito, Toshiaki Yasue, Takeshi Ogasawara, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
OOPSLA8
2002 Bytecode fetch optimization for a Java interpreter
abstract
Interpreters play an important role in many languages, and their performance is critical particularly for the popular language Java. The performance of the interpreter is important even for high-performance virtual machines that employ just-in-time compiler technology, because there are advantages in delaying the start of compilation and in reducing the number of the target methods to be compiled. Many techniques have been proposed to improve the performance of various interpreters, but none of them has fully addressed the issues of minimizing redundant memory accesses and the overhead of indirect branches inherent to interpreters running on superscalar processors. These issues are especially serious for Java because each bytecode is typically one or a few bytes long and the execution routine for each bytecode is also short due to the low-level, stack-based semantics of Java bytecode. In this paper, we describe three novel techniques of our Java bytecode interpreter, write-through top-of-stack caching (WT), position-based handler customization (PHC), and position-based speculative decoding (PSD), which ameliorate these problems for the PowerPC processors. We show how each technique contributes to improving the overall performance of the interpreter for major Java benchmark programs on an IBM POWER3 processor. Among three, PHC is the most effective one. We also show that the main source of memory accesses is due to bytecode fetches and that PHC successfully eliminates the majority of them, while it keeps the instruction cache miss ratios small.
Kazunori Ogata, Hideaki Komatsu, Toshio Nakatani
ASPLOS1