EDBT 2026 Demo / reviewers in the wild / expert
Kiyokuni Kawachiya
dblp:43/3549
· DBLP profile ↗
21ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-authorComputer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 38% Parallel and multicore computing · 22% GPUs and heterogeneous computing · 15% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Software engineering, system software, and programming languages
6 papers |
Runtime systems and virtual machines · 51% Concurrent programming · 14% Operating systems · 12% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.6 | 2 | 2019 | Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019 Resilient X10: efficient failure-aware programming · PPoPP 2014 |
Parallel and multicore computing
parallel programming models |
0.6 | 2 | 2019 | Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019 Resilient X10: efficient failure-aware programming · PPoPP 2014 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.4 | 1 | 2019 | Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019 |
Machine learning › Efficient and distributed learning › memory-efficient training
out-of-core training |
0.4 | 1 | 2019 | Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019 |
Distributed systems › fault tolerance
failure recovery |
0.4 | 1 | 2019 | Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019 |
GPUs and heterogeneous computing
GPU computing |
0.4 | 1 | 2019 | Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019 |
Storage systems
out-of-core computation |
0.4 | 1 | 2019 | Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.1 | 3 | 2010 | A study of Java's non-Java memory · OOPSLA 2010 Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008 Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.1 | 2 | 2006 | Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006 Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003 |
Runtime systems and virtual machines
garbage collection |
0.1 | 1 | 2008 | Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008 |
Operating systems › resource management
memory management |
0.1 | 1 | 2008 | Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008 |
Debugging and program repair › software debugging
compiler debugging |
0.1 | 1 | 2006 | Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006 |
Concurrent programming
synchronization |
0.1 | 2 | 2002 | Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002 A Study of Locking Objects with Bimodal Fields · OOPSLA 1999 |
Program analysis
data flow analysis |
0.0 | 1 | 2003 | Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003 |
Compilers and program optimization › interprocedural optimization
inlining |
0.0 | 1 | 2003 | Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003 |
Concurrent programming › synchronization
locking |
0.0 | 1 | 2002 | Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2010 | A study of Java's non-Java memory · OOPSLA 2010 |
Content delivery and video streaming
video-on-demand |
0.0 | 1 | 1996 | A Portable Communication System for Video-On-Demand Applications Using the Existing Infrastructure · INFOCOM 1996 |
Runtime systems and virtual machines › managed runtime
java runtime |
0.0 | 1 | 1999 | A Study of Locking Objects with Bimodal Fields · OOPSLA 1999 |
Ubiquitous computing and smart environments › mobile computing
mobile web browsing |
0.0 | 1 | 1998 | NaviPoint: An Input Device for Mobile Information Browsing · CHI 1998 |
Transport protocols and congestion control
flow control |
0.0 | 1 | 1996 | A Portable Communication System for Video-On-Demand Applications Using the Existing Infrastructure · INFOCOM 1996 |
Methods — techniques the papers use, named apart from their topics
profiling · 0.8happens-before invariance · 0.4formal semantics · 0.4visualization · 0.2memory statistics gathering · 0.2heap analysis · 0.1garbage collection · 0.1state saving · 0.1replay compilation · 0.1partial redundancy elimination · 0.0exception check elimination · 0.0thread locality exploitation · 0.0user evaluation · 0.0input device design · 0.0user-level library implementation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Acceleration of Large Deep Learning Training with Hybrid GPU Memory Management of Swapping and Re-computingabstractDeep learning has achieved overwhelmingly better accuracy than existing methods in various fields. To further improve deep learning, a deeper and larger neural network (NN) model is indispensable, but graphics processing unit (GPU) memory is not large enough to train such NN models. One promising method to reduce GPU memory consumption is data swapping, which eases the burden on GPU memory by swapping out intermediate data from GPU memory to central processing unit (CPU) memory while the data are not necessary. However, this method introduces communication overhead in transferring data between GPU and CPU memory. Another method is recomputation, which discards intermediate data once and then computes them again when necessary. Unlike the data-swapping method, the re-computation method introduces additional computation but does not require CPU-GPU communication. Therefore, it may reduce CPU-GPU communication by introducing it effectively in the data-swapping method. In this paper, we developed a faster training method for large NN models by combining the re-computation and the data-swapping methods. This method edits the graph of TensorFlow automatically. We developed heuristics about which part of the graph should be recomputed. The heuristics divides the graph into sub-graphs and applies the re-computation in the decreasing order of the amount of swapping data in the sub-graphs. This heuristics is intended to reduce maximal communications by re-computing minimal sub-graphs. Our hybrid method improves performance by up to 15.5% in image size 8000×8000 of ResNet50, 15.3% in image size 7500×7500 of ResNet152, and 12.4% in image size 1000×1000 of DeepLabV3+ compared with the existing data-swapping method. Haruki Imai, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya |
IEEE BigData | 4 |
| 2019 | Automatic GPU memory management for large neural models in TensorFlowabstractDeep learning models are becoming larger and will not fit in the limited memory of accelerators such as GPUs for training. Though many methods have been proposed to solve this problem, they are rather ad-hoc in nature and difficult to extend and integrate with other techniques. In this paper, we tackle the problem in a formal way to provide a strong foundation for supporting large models. We propose a method of formally rewriting the computational graph of a model where swap-out and swap-in operations are inserted to temporarily store intermediate results on CPU memory. By introducing a categorized topological ordering for simulating graph execution, the memory consumption of a model can be easily analyzed by using operation distances in the ordering. As a result, the problem of fitting a large model into a memory-limited accelerator is reduced to the problem of reducing operation distances in a categorized topological ordering. We then show how to formally derive swap-out and swap-in operations from an existing graph and present rules to optimize the graph. Finally, we propose a simulation-based auto-tuning to automatically find suitable graph-rewriting parameters for the best performance. We developed a module in TensorFlow, called LMS, by which we successfully trained ResNet-50 with a 4.9x larger mini-batch size and 3D U-Net with a 5.6x larger image resolution. Tung D. Le, Haruki Imai, Yasushi Negishi, Kiyokuni Kawachiya |
ISMM | 4 |
| 2019 | High Resolution Medical Image Segmentation Using Data-Swapping Method
Haruki Imai, Samuel Matzek, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya |
MICCAI (3) | 5 |
| 2019 | Profiling based out-of-core hybrid method for large neural networks: posterabstractNeural networks (NNs) have archived high accuracy in many fields of machine learning such as image recognition. Since computations of NNs are heavy tasks, GPUs have been widely used to accelerate them. However, the problem sizes of NNs that can be computed are limited by GPU memory capacity. Haruki Imai, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya, Ryo Matsumiya, Toshio Endo |
PPoPP | 5 |
| 2019 | Failure Recovery in Resilient X10abstractCloud computing has made the resources needed to execute large-scale in-memory distributed computations widely available. Specialized programming models, e.g., MapReduce, have emerged to offer transparent fault tolerance and fault recovery for specific computational patterns, but they sacrifice generality. In contrast, the Resilient X10 programming language adds failure containment and failure awareness to a general purpose, distributed programming language. A Resilient X10 application spans over a number of places. Its formal semantics precisely specify how it continues executing after a place failure. Thanks to failure awareness, the X10 programmer can in principle build redundancy into an application to recover from failures. In practice, however, correctness is elusive, as redundancy and recovery are often complex programming tasks. This article further develops Resilient X10 to shift the focus from failure awareness to failure recovery, from both a theoretical and a practical standpoint. We rigorously define the distinction between recoverable and catastrophic failures. We revisit the happens-before invariance principle and its implementation. We shift most of the burden of redundancy and recovery from the programmer to the runtime system and standard library. We make it easy to protect critical data from failure using resilient stores and harness elasticity—dynamic place creation—to persist not just the data but also its spatial distribution. We demonstrate the flexibility and practical usefulness of Resilient X10 by building several representative high-performance in-memory parallel application kernels and frameworks. These codes are 10× to 25× larger than previous Resilient X10 benchmarks. For each application kernel, the average runtime overhead of resiliency is less than 7%. By comparing application kernels written in the Resilient X10 and Spark programming models, we demonstrate that Resilient X10’s more general programming model can enable significantly better application performance for resilient in-memory distributed computations. David Grove, Sara S. Hamouda, Benjamin Herta, Arun Iyengar, Kiyokuni Kawachiya, Josh Milthorpe, Vijay A. Saraswat, Avraham Shinnar, Mikio Takeuchi, Olivier Tardieu |
ACM Trans. Program. Lang. Syst. | 5 |
| 2018 | Involving CPUs into Multi-GPU Deep LearningabstractThe most important part of deep learning, training the neural network, often requires the processing of a large amount of data and can takes days to complete. Data parallelism is widely used for training deep neural networks on multiple GPUs in a single machine thanks to its simplicity. However, its scalability is bound by the number of data transfers, mainly for exchanging and accumulating gradients among the GPUs. In this paper, we present a novel approach to data parallel training called CPU-GPU data parallel (CGDP) training that utilizes free CPU time on the host to speed up the training in the GPUs. We also present a cost model for analyzing and comparing the performances of both the typical data parallel training and the CPU-GPU data parallel training. Using the cost model, we formally show why our approach is better than the typical one and clarify the remaining issues. Finally, we explain how we optimized CPU-GPU data parallel training by introducing chunks of layers and present a runtime algorithm that automatically finds a good configuration for the training. The algorithm is effective for very deep neural networks, which are the current trend in deep learning. Experimental results showed that we achieved speedups of $1.21$, $1.04$, $1.21$ and $1.07$ for four state-of-the-art neural networks: AlexNet, GoogLeNet-v1, VGGNet-16, and ResNet-152, respectively. Weak scaling efficiency greater than $90$ was achieved for all networks across four GPUs. Tung D. Le, Taro Sekiyama, Yasushi Negishi, Haruki Imai, Kiyokuni Kawachiya |
ICPE | 5 |
| 2014 | Resilient X10: efficient failure-aware programmingabstractScale-out programs run on multiple processes in a cluster. In scale-out systems, processes can fail. Computations using traditional libraries such as MPI fail when any component process fails. The advent of Map Reduce, Resilient Data Sets and MillWheel has shown dramatic improvements in productivity are possible when a high-level programming framework handles scale-out and resilience automatically. David Cunningham, David Grove, Benjamin Herta, Arun Iyengar, Kiyokuni Kawachiya, Hiroki Murata, Vijay A. Saraswat, Mikio Takeuchi, Olivier Tardieu |
PPoPP | 5 |
| 2014 | String deduplication for Java-based middleware in virtualized environmentsabstractTo increase the memory efficiency in physical servers is a significant concern for increasing the number of virtual machines (VM) in them. When similar web application service runs in each guest VM, many string data with the same values are created in every guest VMs. These duplications of string data are redundant from the viewpoint of memory efficiency in the host OS. This paper proposes two approaches to reduce the duplication in Java string in a single Java VM (JVM) and across JVMs. The first approach is to share string objects cross JVMs by using a read-only memory-mapped file. The other approach is to selectively unify string objects created at runtime in the web applications. This paper evaluates our approach by using the Apache DayTrader and the DaCapo benchmark suite. Our prototype implementation chieved 7% to 12% reduction in the total size of the objects allocated over the lifetime of the programs. In addition, we observed the performance of DayTrader was maintained even under a situation of high density guest VMs in a KVM host machine. Michihiro Horie, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera |
VEE | 3 |
| 2010 | A study of Java's non-Java memoryabstractA Java application sometimes raises an out-of-memory ex-ception. This is usually because it has exhausted the Java heap. However, a Java application can raise an out-of-memory exception when it exhausts the memory used by Java that is not in the Java heap. We call this area non-Java memory. For example, an out-of-memory exception in the non-Java memory can happen when the JVM attempts to load too many classes. Although it is relatively rare to ex-haust the non-Java memory compared to exhausting the Java heap, a Java application can consume a considerable amount of non-Java memory.This paper presents a quantitative analysis of non-Java memory. To the best of our knowledge, this is the first in-depth analysis of the non-Java memory. To do this we cre-ated a tool called Memory Analyzer for Redundant, Unused, and String Areas (MARUSA), which gathers memory statis-tics from both the OS and the Java virtual machine, break-ing down and visualizing the non-Java memory usage.We studied the use of non-Java memory for a wide range of Java applications, including the DaCapo benchmarks and Apache DayTrader. Our study is based on the IBM J9 Java Virtual Machine for Linux. Although some of our results may be specific to this combination, we believe that most of our observations are applicable to other platforms as well. Kazunori Ogata, Dai Mikurube, Kiyokuni Kawachiya, Scott Trent, Tamiya Onodera |
OOPSLA | 3 |
| 2010 | Efficient runtime tracking of allocation sites in JavaabstractTracking the allocation site of every object at runtime is useful for reliable, optimized Java. To be used in production environments, the tracking must be accurate with minimal speed loss. Previous approaches suffer from performance degradation due to the additional field added to each object or track the allocation sites only probabilistically. We propose two novel approaches to track the allocation sites of every object in Java with only a 1.0% slow-down on average. Our first approach, the Allocation-Site-as-a-Hash-code (ASH) Tracker, encodes the allocation site ID of an object into the hash code field of its header by regarding the ID as part of the hash code. ASH Tracker avoids an excessive increase in hash code collisions by dynamically shrinking the bit-length of the ID as more and more objects are allocated at that site. For those Java VMs without the hash code field, our second approach, the Allocation-Site-via-a-Class-pointer (ASC) Tracker, makes the class pointer field in an object header refer to the allocation site structure of the object, which in turn points to the actual class structure. ASC Tracker mitigates the indirection overhead by constant-class-field duplication and allocation-site equality checks. While a previous approach of adding a 4-byte field caused up to 14.4% and an average 5% slowdown, both ASH and ASC Trackers incur at most a 2.0% and an average 1.0% loss. We demonstrate the usefulness of our low-overhead trackers by an allocation-site-aware memory leak detector and allocation-site-based pretenuring in generational GC. Our pretenuring achieved on average 1.8% and up to 11.8% speedups in SPECjvm2008. Rei Odaira, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera, Toshio Nakatani |
VEE | 3 |
| 2008 | Analysis and reduction of memory inefficiencies in Java stringsabstractThis paper describes a novel approach to reduce the memory consumption of Java programs, by focusing on their "string memory inefficiencies". In recent Java applications, string data occupies a large amount of the heap area. For example, about 40% of the live heap area is used for string data when a production J2EE application server is running. By investigating the string data in the live heap, we identified two types of memory inefficiencies -- "duplication" and "unused literals". In the heap, there are many string objects that have the same values. There also exist many string literals whose values are not actually used by the application. Since these inefficiencies exist as live objects, they cannot be eliminated by existing garbage collection techniques, which only remove dead objects. Quantitative analysis of Java heaps in real applications revealed that more than 50% of the string data in the live heap is wasted by these inefficiencies. To reduce the string memory inefficiencies, this paper proposes two techniques at the Java virtual machine level, "StringGC" for eliminating duplicated strings at the time of garbage collection, and "Lazy Body Creation" for delaying part of the literal instantiation until the literal's value is actually used. We also present an interesting technique at the Java program level, which we call "BundleConverter", for preventing unused message literals from being instantiated. Prototype implementations on a production Java virtual machine have achieved about 18% reduction of the live heap in the production application server. The proposed techniques could also reduce the live heap of standard Java benchmarks by 11.6% on average, without noticeable performance degradation. Kiyokuni Kawachiya, Kazunori Ogata, Tamiya Onodera |
OOPSLA | 1 |
| 2007 | Libra: a library operating system for a jvm in a virtualized execution environmentabstractIf the operating system could be specialized for every application, many applications would run faster. For example, Java virtual machines (JVMs) provide their own threading model and memory protection, so general-purpose operating system implementations of these abstractions are redundant. However, traditional means of transforming existing systems into specialized systems are difficult to adopt because they require replacing the entire operating system. This paper describes Libra, an execution environment specialized for IBM's J9 JVM. Libra does not replace the entire operating system. Instead, Libra and J9 form a single statically-linked image that runs in a hypervisor partition. Libra provides the services necessary to achieve good performance for the Java workloads of interest but relies on an instance of Linux in another hypervisor partition to provide a networking stack, a filesystem, and other services. The expense of remote calls is offset by the fact that Libra's services can be customized for a particular workload; for example, on the Nutch search engine, we show that two simple customizations improve application throughput by a factor of 2.7. Glenn Ammons, Jonathan Appavoo, Maria A. Butrico, Dilma Da Silva, David Grove, Kiyokuni Kawachiya, Orran Krieger, Bryan S. Rosenburg, Eric Van Hensbergen, Robert W. Wisniewski |
VEE | 6 |
| 2007 | Cloneable JVM: a new approach to start isolated java applications fasterabstractJava has been successful particularly for writing applications in the server environment. However, isolation of multiple applications hasnot been efficiently achieved in Java. Many customers require that their applications are guarded by independent OS processes, but starting a Java application with a new process results in a long sequence of initializations being repeated each time. To date, there has been no way to quickly start a new Java application as an isolated OS process. In this paper, we propose a new isolation approach called Cloneable JVM to eliminate this startup overhead in Java. The key idea is to createa new Java application by copying, or cloning, the already-initialized image of the primary JVM process. Since the clone is already initialized, it can begin actual operations immediately as a new isolated process. This cloning abstraction can support new scenarios for Java, such as user isolation and transaction isolation. We implemented a prototype of the Cloneable JVM by modifying a production JVM on Linux, which provides a new API for cloning constructed on the Isolate API defined in JSR 121. Using this cloning API, several Java applications, including a large production J2EE application server, we remodified to demonstrate the isolation scenarios. Evaluations using these prototypes showed that new ready-to-serve Java applications can start up as a new process in less than 5 seconds, which is 4 to 170 times faster than starting these applications from scratch. Kiyokuni Kawachiya, Kazunori Ogata, Daniel Silva 0001, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani |
VEE | 1 |
| 2006 | Replay compilation: improving debuggability of a just-in-time compilerabstractThe performance of Java has been tremendously improved by the advance of Just-in-Time (JIT) compilation technologies. However, debugging such a dynamic compiler is much harder than a static compiler. Recompiling the problematic method to produce a diagnostic output does not necessarily work as expected, because the compilation of a method depends on runtime information at the time of compilation.In this paper, we propose a new approach, called replay JIT compilation, which can reproduce the same compilation remotely by using two compilers, the state-saving compiler and the replaying compiler. The state-saving compiler is used in a normal run, and, while compiling a method, records into a log all of the input for the compiler. The replaying compiler is then used in a debugging run with the system dump, to recompile a method with the options for diagnostic output. We reduced the overhead to save the input by using the system dump and by categorizing the input based on how its value changes. In our experiment, the increase of the compilation time for saving the input was only 1%, and the size of the additional memory needed for saving the input was only 10% of the compiler-generated code. Kazunori Ogata, Tamiya Onodera, Kiyokuni Kawachiya, Hideaki Komatsu, Toshio Nakatani |
OOPSLA | 3 |
| 2004 | Lock Reservation for Java Reconsidered
Tamiya Onodera, Kiyokuni Kawachiya, Akira Koseki |
ECOOP | 2 |
| 2003 | Effectiveness of cross-platform optimizations for a java just-in-time compilerabstractThis paper describes the system overview of our Java Just-In-Time (JIT) compiler, which is the basis for the latest production version of IBM Java JIT compiler that supports a diversity of processor architectures including both 32-bit and 64-bit modes, CISC, RISC, and VLIW architectures. In particular, we focus on the design and evaluation of the cross-platform optimizations that are common across different architectures. We studied the effectiveness of each optimization by selectively disabling it in our JIT compiler on three different platforms: IA-32, IA-64, and PowerPC. Our detailed measurements allowed us to rank the optimizations in terms of the greatest performance improvements with the smallest compilation times. The identified set includes method inlining only for tiny methods, exception check eliminations using forward dataflow analysis and partial redundancy elimination, scalar replacement for instance and class fields using dataflow analysis, optimizations for type inclusion checks, and the elimination of merge points in the control flow graphs. These optimizations can achieve 90% of the peak performance for two industry-standard benchmark programs on these platforms with only 34% of the compilation time compared to the case for using all of the optimizations. Kazuaki Ishizaki, Mikio Takeuchi, Kiyokuni Kawachiya, Toshio Suganuma, Osamu Gohda, Tatsushi Inagaki, Akira Koseki, Kazunori Ogata, Motohiro Kawahito, Toshiaki Yasue, Takeshi Ogasawara, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani |
OOPSLA | 3 |
| 2002 | Lock reservation: Java locks can mostly do without atomic operationsabstractBecause of the built-in support for multi-threaded programming, Java programs perform many lock operations. Although the overhead has been significantly reduced in the recent virtual machines, One or more atomic operations are required for acquiring and releasing an object's lock even in the fastest cases.This paper presents a novel algorithm called lock reservation. It exploits thread locality of Java locks, which claims that the locking sequence of a Java lock contains a very long repetition of a specific thread. The algorithm allows locks to be reserved for threads. When a thread attempts to acquire a lock, it can do without any atomic operation if the lock is reserved for the thread. Otherwise, it cancels the reservation and falls back to a conventional locking algorithm.We have evaluated an implementation of lock reservation in IBM's production virtual machine and compiler. The results show that it achieved performance improvements up to 53% in real Java programs. Kiyokuni Kawachiya, Akira Koseki, Tamiya Onodera |
OOPSLA | 1 |
| 1999 | A Study of Locking Objects with Bimodal FieldsabstractObject locking can be efficiently implemented by bimodal use of a field reserved in an object. The field is used as a lightweight lock in one mode, while it holds a reference to a heavyweight lock in the other mode. A bimodal locking algorithm recently proposed for Java achieves the highest performance in the absence of contention, and is still fast enough when contention occurs. Tamiya Onodera, Kiyokuni Kawachiya |
OOPSLA | 2 |
| 1998 | NaviPoint: An Input Device for Mobile Information BrowsingabstractArticle Free Access Share on NaviPoint: an input device for mobile information browsing Authors: Kiyokuni Kawachiya IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, Japan IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, JapanView Profile , Hiroshi Ishikawa IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, Japan IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, JapanView Profile Authors Info & Claims CHI '98: Proceedings of the SIGCHI Conference on Human Factors in Computing SystemsJanuary 1998 Pages 1–8https://doi.org/10.1145/274644.274645Published:01 January 1998Publication History 14citation985DownloadsMetricsTotal Citations14Total Downloads985Last 12 Months13Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Kiyokuni Kawachiya, Hiroshi Ishikawa 0008 |
CHI | 1 |
| 1996 | A Portable Communication System for Video-On-Demand Applications Using the Existing InfrastructureabstractVideo on demand (VOD) systems are becoming popular. They have special requirements for receiving VOD data steadily, but common general communication protocols cannot meet their requirements. Several new protocols have been designed to fulfil this service requirement, but it takes a long time for a protocol to be supported from end to end. We propose a VOD communication system that is implemented as an user-level library on top of a common existing protocol. This protocol is easy to port, and the level of connectivity is almost the same as in the underlying existing protocol. We describe the important points related to its implementation, such as the low overhead of implementation and the flow control mechanism. We implemented a prototype system, measured its performance, and concluded that implementing a communication system as a user-level library enables us to implement a VOD portable communication system with a practical level of performance. Yasushi Negishi, Kiyokuni Kawachiya, Kazuya Tago |
INFOCOM | 2 |
| 1995 | Evaluation of QoS-Control Servers on Real-Time Mach
Kiyokuni Kawachiya, Masanobu Ogata, Nobuhiko Nishio, Hideyuki Tokuda |
NOSSDAV | 1 |